VLDB 2026 Research / reviewers in the wild / expert
Yuwen He
dblp:35/1286
· DBLP profile ↗
51ranked-venue papers
10as first author
13since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 44 · 9 first-author · 9 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 1 since 2021Systems, architecture and hardware · 5 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Search Method for Approximate Optimal Rate Control Solution via Reinforcement LearningabstractRecent studies on video rate control (RC) have introduced accurate and high performance methods but have not explored the optimal RC solution. The optimal solution is crucial for improving RC methods and providing labels for supervised learning. To find approximate optimal RC solutions within a limited time, we are the first to propose a reinforcement learning based search method for finding approximate optimal RC solutions within a limited time. Specifically, the RC problem for a video is first modeled as a Markov decision process (MDP). Then, with the MDP model, we develop a search method based on the deep Q-network method, which consists of exploration and exploitation steps. During exploration, an agent is created, consisting of two multilayer perceptrons and a replay memory, and trained within the MDP to estimate the value function, while superior RC solutions are recorded throughout the training process. After training, RC solution is estimated by the trained agent using the value function and a greedy strategy during the exploitation step. Finally, the approximate optimal RC solution is determined based on the RC solutions from both two steps. In addition, the time complexity of proposed method is controllable, specifically,$O(m n)$where$m$denotes the number of training epoch. Experimental results show that the bit-rate error and compression quality of the solutions found by proposed method approach the optimal solutions, with differences of only less than 0.005% and 0.399%, respectively, and are achieved in a significantly shorter time compared to the brute force search. Longtao Feng, Qian Yin 0002, Jiaqi Zhang 0007, Yuwen He, Siwei Ma 0001 |
DCC | 5 |
| 2026 | High Accuracy Rate Control for Neural Video Coding Based on Rate-Distortion ModelingabstractIn recent years, rate control (RC) for neural video coding (NVC) has become an active research area. However, existing RC methods in NVC neglect the actual rate-distortion (R-D) characteristics and lack dedicated optimization strategies for intra and inter modes, leading to significant bit rate errors. To address these issues, we propose a high accuracy RC method for NVC based onR-Dmodeling, which integrates intra frame RC, inter frame RC and bit allocation. Specifically, the rate-quantization parameter (R-Q) model andR-Dmodel are established for both intra frame and inter frame in NVC. To derive the model parameters, intra frame parameters are estimated using high dimensional features, while inter frame parameters are derived using gradient descent based model update methods. Based on the proposedR-Qmodel, intra frame and inter frame RC methods are proposed to determine the quantization parameters (QP). Meanwhile, a bit allocation method is developed based on the derivedR-Dmodels to allocate bits for the intra frame and inter frame. Extensive experiments demonstrate that, benefiting from the accurateR-Qmodels derived by the proposed approach, highly accurate RC is achieved with only 0.56% average bit rate error. Compared with other methods, the proposed method reduces the average bit rate error by more than 4.18%, and achieves over 8.94% Bjøntegaard Delta Rate savings. Longtao Feng, Qian Yin 0002, Jiaqi Zhang 0007, Yuwen He, Siwei Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Disentangle Nighttime Lens Flares: Self-supervised Generation-based Lens Flare RemovalabstractLens flares arise from light reflection and refraction within sensor arrays, whose diverse types include glow, veiling glare, reflective flare and so on. Existing methods are specialized for one specific type only, and overlook the simultaneous occurrence of multiple typed lens flares, which is common in the real-world, e.g. coexistence of glow and displacement reflections from the same light source. These co-occurring lens flares cannot be effectively resolved by the simple combination of individual flare removal methods, since these coexisting flares originates from the same light source and are generated simultaneously within the same sensor array, exhibit a complex interdependence rather than simple additive relation. To model this interdependent flares’ relationship, our Nighttime Lens Flare Formation model is the first attempt to learn the intrinsic physical relationship between flares on the imaging plane. Building on this physical model, we introduce a solution to this joint flare removal task named Self-supervised Generation-based Lens Flare Removal Network (SGLFR-Net), which is self-supervised without pre-training. Specifically, the nighttime glow is detangled in PSF Rendering Network(PSFR-Net) based on PSF Rendering Prior, while the reflective flare is modelled in Texture Prior Based Reflection Flare Removal Network (TPRR-Net). Empirical evaluations demonstrate the effectiveness of the proposed method in both joint and individual glare removal tasks. Yuwen He, Wanyu Wu, Kui Jiang |
AAAI | 1 |
| 2025 | Adaptive Upscaling Filter Selection with Wiener Filter Compression Techniques in Reference Picture ResamplingabstractEncoding video frames at adaptive resolutions has been shown to be an effective strategy, as it balances the trade-off between side information and residual data based on frame size and content. Reference Picture Resampling (RPR) in the Versatile Video Coding (H.266/VVC) standard allows for motion compensation between frames with different resolutions, overcoming the limitation that resolution switching points must occur at intra random-access points (IRAP), and that all frames within a closed GOP must be encoded at a single resolution. Frame rescaling is a critical step in RPR; however, applying a single upscaling filter to all video frames of different resolutions is not ideal. In this paper, an adaptive upscaling filter selection scheme is proposed, which incorporates Wiener filter trained at the frame level. The rate-distortion (RD) performance of various upscaling filters is compared to determine the optimal adaptive filter. The Wiener filter coefficients are reduced and compressed using predictive coding. Experimental results on VVC test model (VTM) 23.3 show average BD rate gains of 0.33% for 4K, 2.15% for 1080p and 2.11% for 480p video sequences at a 2.0x RPR scaling ratio under random-access configurations. At a 1.5x RPR scaling ratio, the BD rate gains are 0.57% for 4K, 2.22% for 1080p and 4.71% for 480p video sequences. Yuwen He, Kenneth Rose |
ISCAS | 3 |
| 2024 | Spatial Neighbor Information Assisted Motion Compensated Temporal Filter for Video CodingabstractMotion compensated temporal filter (MCTF) is a pre-filtering technology for video encoding, which employs bilateral filtering to enhance temporal correlations among adjacent frames. In this paper, a spatial neighbor information assisted MCTF method is proposed, which introduces the spatial information into the motion estimation and bilateral filtering processes in MCTF to improve the coding efficiency. For the motion estimation process, the spatial neighbor information is involved in preventing motion estimation from capturing a local optimum. For the bilateral filtering process, the spatial neighbor information is employed as well to reduce the block boundary effect. The proposed method is implemented on top of the Fraunhofer Versatile Video Encoder (VVenC). Experimental results show that the proposed method can achieve 1.05% bitrate saving on average compared to the current MCTF scheme, and 10.33% bitrate saving compared with MCTF disabled. Zikun Yuan, Weijia Zhu, Yuwen He, Xiaohu Tang 0004 |
PCS | 3 |
| 2024 | Adaptive Block-Level Quality Parameter Adjustment Towards Low Video Bit-Rate FluctuationabstractExisting quantization parameter (QP) adjustment methods in video coding often focus solely on coding efficiency and ignore the impact of bit-rate fluctuations on video transmission and bandwidth waste. This is mainly because intra pictures, in a hierarchical coding structure, are allocated smaller QP and thus consume more bits. To address this issue, we propose an adaptive block-level QP adjustment method. Specifically, intra picture importance (IPI) is first introduced to evaluate the adjustability of intra picture QP. For intra pictures whose QP can be adjusted, we further propose block importance (BI) to determine their optimal block-level QP adjustment. Experimental results show that our proposed method reduce the bit-rate fluctuations while basically maintaining the coding performance. Notably, significant improvements can be observed in high-resolution videos, with a reduction of approximately 11% in bit-rate fluctuations. Longtao Feng, Qian Yin 0002, Huiwen Ren, Zhao Wang 0004, Siwei Ma 0001, Yuwen He |
VCIP | 6 |
| 2023 | High-Precision Motion Vector Refinement for Bi-Directional Optical FlowabstractBi-directional Optical Flow (BDOF), is a very effective tool, developed based on the optical flow concept, which assumes that the motion of an object is smooth and is in a straight line in a short period of time. BDOF as a prediction adjustment tool for each sample, is included in VVC SW, where it derives parameters for each 4×4 block and adjusts each of its samples individually. The reference SW which is being developed for beyond VVC activity, i.e. ECM, has 2 BDOFs: One as Motion Vector (MV) refinement tool for each 8 × 8 subblock, and another one as sample adjustment BDOF similar to VVC one, but deriving the parameters for each sample. In this paper, a High-Precision BDOF is introduced which consists of the following parts: Firstly, a more accurate BDOF formula is used to derive the parameters for BDOF MV refinement. Next, a position-dependent weight is used to increase the focus of the parameter derivation. Finally, higher granularity as small as 4×4 is used for each subblock. For BDOF sample adjustment, another position-dependent weighted sum is used. The proposed method is implemented on top of ECM-7.0. Experimental results show that the proposed method provides up to –0.92% bitrate saving for some sequences compared to the latest ECM SW under random access configuration. The method explained in this paper is adopted into ECM-9.0. Mehdi Salehifar, Yuwen He, Kai Zhang 0007, Li Zhang 0006 |
ICIP | 2 |
| 2023 | Iterative Bi-Directional Optical Flow for Decoder Side Motion Vector RefinementabstractMulti-pass Decoder Side Motion Vector Refinement (DMVR) based on a Bilateral Matching (BM) concept is a very important tool being used in video codecs to improve the Motion Vector (MV) accuracy on the decoder side. In VVC, a single stage DMVR is applied on each 16x16 subblock to refine its MV. Exploring Coding Model (ECM), which is being developed for beyond VVC activity, adopts multi-pass DMVR. First, a BM approach is used to adjust the MV for each Prediction Unit (PU). Next, another BM approach is used to adjust the MVs for each 16x16 subblock. Finally, a Bi-directional Optical Flow (BDOF) based approach is used to adjust the MVs for each 8x8 subblock. In this paper, an iterative BDOF to improve multi-pass DMVR is proposed, where an additional pass of BDOF is added as the 4th stage of the multi-pass DMVR. To balance the gain and the complexity, the granularity of subblocks is adjusted adaptively in the third and forth stages of multi-pass DMVR. Finally, a regularization factor is added to avoid unnecessary over-tuning of BDOF adjustment. The proposed method is implemented on top of ECM-8.0. Experimental results show that the proposed method provides more than -1% bitrate saving for some sequences compared to ECM-8.0 under random access configurations. This method is adopted into ECM-10.0. Mehdi Salehifar, Yuwen He, Kai Zhang 0007, Hongbin Liu 0004, Li Zhang 0006 |
VCIP | 2 |
| 2022 | Learning-Based End-to-End Video Compression with Spatial-Temporal AdaptationabstractThe learning-based end-to-end video compression exhibits a fast development with continuous improvements. In previous works, a key frame in random-access scenarios is typically compressed by an image compressor, and the remaining frames are reconstructed by interpolations. But solely using image compression fails to leverage the temporal correlations, which is critical to achieve a substantial gain in video compression. To exploit temporal correlations among key frames, we introduce a learning-based end-to-end spatial-temporal adaptive (e2e-STA) compression solution to offer flexible options for the key frames. First of all, we design an extrapolation-based key frame compression scheme. Given a key frame, e2e-STA can switch between an image compressor and an extrapolative compressor. A key frame is thereby able to select the optimal solution adaptively according to the rate-distortion optimization criteria and the optimal selection is sent to the decoder. The proposed approach is optimized end-to-end with all the networks. The experimental results validate the effectiveness of the proposed mechanism. The proposed method outperforms existing learning-based video compression methods by a noticeable margin and provides promising performance compared to traditional benchmark video codecs in terms of PSNR and MS-SSIM. Zhaobin Zhang, Yue Li 0015, Kai Zhang 0007, Li Zhang 0006, Yuwen He |
ICIP | 5 |
| 2022 | A Template Matching based Extension for Merge with Motion Vector DifferenceabstractMerge with motion vector difference (MMVD) is a simple yet efficient coding tool adopted in VVC. However, the efficiency of MMVD drops a lot in the reference software beyond VVC, known as ECM, this paper proposes a template matching (TM) based extension for MMVD (TME-MMVD) for inter prediction. Firstly, extra refinement positions are introduced along diagonal angles. Secondly, based on a template matching cost calculated as a distortion between samples of a template and their reference samples, the refinement positions are reordered. Finally, the preferable ones with the lowest costs are selected as available positions and consequently for index coding. This method is also extended to the affine MMVD mode. The proposed method is implemented on top of the ECM2. Experimental results show that the proposed method provides 0.69% and 0.64% bitrate saving compared to latest VVC SW and 0.24% and 0.31% bitrate saving compared to ECM2 under the JVET common test conditions for random access and low delay configurations, respectively. This tool is adopted into the current ECM SW. Mehdi Salehifar, Yuwen He, Kai Zhang 0007, Li Zhang 0006 |
ISCAS | 2 |
| 2022 | Optimized Bit Allocation for Learning-based Video CompressionabstractThe optimized bit allocation among frames has been intensively explored and improved the compression performance significantly in conventional video coding. However, the optimized bit allocation is still in its infant stage for learning-based video coding. Most existing learning-based video compression methods either use uniform bit allocation or empirically determined bit allocation weights among frames. In this paper, we develop an optimized bit allocation scheme for learning-based end-to-end video compression. In particular, we realize a hierarchical quality-control mechanism based on the importance of different frames under random-access scenarios. Considering the varying importance of frames on different temporal layers, we propose an efficient yet simple scheme, in which a set of optimized bit allocation weights are introduced to the rate-distortion (R-D) loss function. Experimental results demonstrate the effectiveness of the proposed scheme. In addition, the proposed scheme can be easily applied to most existing learning-based video compression frameworks under random-access scenarios. Zhaobin Zhang, Yue Li 0015, Kai Zhang 0007, Li Zhang 0006, Yuwen He |
ISCAS | 5 |
| 2021 | Probability-based decoder-side intra mode derivation for VVCabstractIntra prediction is typically used to exploit the spatial redundancy in video coding. In the latest video coding standard Versatile Video Coding (VVC), 67 intra prediction modes are adopted in intra prediction. The encoder selects the best one from 67 modes and signals it to the decoder. Bits consuming of signaling the selected mode may limit the coding efficiency. To reduce the overhead of signaling the intra prediction mode, a probability-based decoder-side intra mode derivation (P-DIMD) is proposed in this paper. Specifically, an intra prediction mode candidate set is constructed based on the probabilities of intra prediction modes. The probability of an intra prediction mode is mainly estimated in two ways. First, the textures are typically continuous within a local region and intra prediction modes of neighboring blocks are similar to each other. Second, some intra prediction modes are preferable to be used than others. For each intra prediction mode in the constructed candidate set, intra prediction is processed on a template to calculate a cost. The intra prediction mode with the minimum cost is determined as the optimal mode and used in the intra prediction of the current block. Experimental results demonstrate that P-DIMD can achieve 0.56% BD-rate saving on average compared to VTM-11.0 under all intra configuration. Li Zhang 0006, Kai Zhang 0007, Yuwen He, Hongbin Liu 0004 |
VCIP | 4 |
| 2021 | Comparative viromes of Culicoides and mosquitoes reveal their consistency and diversity in viral profilesabstractThe genus Culicoides includes biting midges, some of which are vectors for viruses that cause diseases in humans and animals. Knowledge of the roles of Culicoides in viral ecology is inadequate. We collected ~300 000 samples of Culicoides and mosquitoes in 15 representative regions within Yunnan, China. Using mosquitoes as reference vectors, we designed a comparative virome strategy to study the viral composition, diversity, hosts and spatiotemporal distribution of Culicoides. A map of viromes in Culicoides and mosquitoes in Yunan province, China, was constructed. At the same locations, Culicoides and mosquitoes usually share a similar viral diversity. At least 10 important pathogenic viruses were detected from Culicoides. Many novel viruses were discovered, including 21 segmented viruses of Flaviviridae, 180 viruses of Monjiviricetes and 130 viruses of Bunyavirales. The findings demonstrate that Culicoides is an important part of viral ecology and should be studied and monitored for potentially emerging viruses. Qin Shen, Yuwen He, Na Han, Xianyue Wang, Jinxin Meng, Yousong Peng, Mei Pan, Yuting Jin, Taijiao Jiang, Wenjie Tan, Jinglin Wang, Aiping Wu 0002 |
Briefings Bioinform. | 4 |
| 2020 | Local Context Attention for Salient Object Segmentation
Pengfei Xiong, Zhengyi Lv, Kuntao Xiao, Yuwen He |
ACCV (1) | 5 |
| 2020 | A Unified Video Codec for SDR, HDR, and 360° Video ApplicationsabstractThe ITU-T Video Coding Experts Group (VCEG) and ISO/IEC Moving Picture Experts Group (MPEG) issued in October 2017 a joint Call for Proposals (CfP) on video compression with capability beyond HEVC. The joint CfP included three categories of content: standard dynamic range (SDR), high dynamic range and wide color gamut (HDR/WCG), and 360° omni-directional video (360°). This paper describes a response to the joint CfP that considers all three categories of video content. The core codec in the response is designed based on the joint exploration model (JEM) reference software. The key coding tools in the JEM are significantly simplified to reduce both average and worst-case complexity for hardware design with negligible coding performance loss. Furthermore, two additional coding tools are used to further improve coding efficiency. For the HDR and 360° categories, additional coding tools specifically designed to optimize the compression efficiency and subjective quality of that specific content category are included. Further, some SDR coding tools are modified to alleviate subjective quality problems. For the random access configuration, compared to the HEVC test model (HM) anchor, the proposed video codec achieves average luma rate savings of 35.7%, 31.3%, and 33.9% for the SDR, HDR, and 360° categories, respectively. Xiaoyu Xiu, Philippe Hanhart, Yuwen He, Yan Ye 0003, Rahul Vanam, Taoran Lu, Fangjun Pu, Peng Yin 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Improved Video Coding Techniques for Next Generation Video Coding StandardabstractThis paper describes a video coding scheme submitted in response to the joint call for proposals (CfP) on video compression for capability beyond HEVC issued by ITU-T SG16 Q.6 (VCEG) and ISO/IEC JTC1/SC29/WG11(MPEG) in October 2017. It includes video coding techniques for the standard dynamic range (SDR) and high dynamic range (HDR) categories. Design of the core SDR codec in the response is based on the joint exploration model (JEM) reference software developed by the joint video exploration team (JVET). Some of key coding tools in the JEM are significantly simplified to reduce both average and worst-case complexity for hardware design with negligible coding performance loss. Furthermore, two additional coding technologies, namely multi-type tree (MTT) and decoder-side intra mode derivation (DIMD), are used to further improve coding efficiency. For the HDR category, besides the tools used in SDR category, two additional coding tools: an in-loop reshaper and a luma-based QP prediction method are used to further improve HDR coding efficiency. Simulation results demonstrate the high coding efficiency achieved by the proposed video codec at the expense of moderate coding complexity over HEVC. For random access configuration, it achieves average bit rate savings of 35.7% and 4.00% over the HM and JEM anchors with decoding time of 263% and 33%, respectively, for the SDR sequences. For the HDR sequences, the proposed in-loop reshaper is configured to maximize HDR objective metrics, it achieves average bit rate savings of 31.3% and 4.6% over the HM and JEM for wPSNRY metrics for the HDR PQ content. Xiaoyu Xiu, Yuwen He, Yan Ye 0003, Rahul Vanam, Philippe Hanhart, Taoran Lu, Fangjun Pu, Peng Yin 0002, Walt Husak, Tao Chen 0044 |
DCC | 2 |
| 2019 | Adaptive Motion Vector Resolution for Affine-Inter Mode CodingabstractAffine Motion Model (AMM) based inter prediction, which can represent complex motions such as zooming, rotation or shearing, has been adopted into the Versatile Video Coding (VVC) standard. AMM is defined by Control Point Motion Vectors (CPMVs) in VVC. On the other hand, Adaptive Motion Vector Resolution (AMVR) has also been adopted into VVC standard due to a favorable trade-off between the Motion Vector (MV) precision and the bit consumption on MV Differences (MVDs). However, AMVR is only applied to the Translational Motion Model (TMM), and AMM cannot benefit from it. In this paper, it is proposed to extend AMVR to AMM. Specifically, 1-pixel, 1/4-pixel and 1/16-pixel MV precisions are allowed and can be selected adaptively by each affine-inter mode coded Coding Unit (CU). Simulation results reportedly show that the proposed method can achieve 0.32% BD-rate saving on average under the Random Access configuration. Hongbin Liu 0004, Li Zhang 0006, Kai Zhang 0007, Jizheng Xu, Yue Wang 0032, Jiancong Luo, Yuwen He |
PCS | 7 |
| 2019 | Prediction Refinement with Optical Flow for Affine Motion CompensationabstractAffine motion compensation (AMC) has been adopted in the latest working draft of the Versatile Video Coding (VVC) standard jointly developed by ITU-T VCEG and ISO/IEC MPEG. The AMC in the VVC working draft is implemented as a sub-block based MC rather than pixel-based MC in order to reduce the memory access bandwidth and computation complexity, which loses prediction efficiency. The proposed algorithm in this paper is to improve the MC granularity by using pixel-based optical flow refinement. The affine prediction efficiency will be close to the pixel-based MC refinement without increasing external memory access. The experiment results show that on average the proposed algorithm can achieve 0.86% and 0.87% BD rate saving for random access and low delay B configurations, respectively, comparing to the VTM-4.0. The BD rate saving is up to 3.10%. Jiancong Luo, Yuwen He |
VCIP | 2 |
| 2018 | Hybrid Cubemap Projection Format for 360-Degree Video Codingabstract360-degree video has become popular in recent years with the advances in virtual reality (VR) and augmented reality (AR) technologies and has been rapidly commercialized. To provide viewers with an immersive experience, 360-degree video requires higher resolution and much higher bandwidth compared with conventional 2D video. In a typical 360-degree video compression and delivery framework, the stitched input 360-degree videos, represented in a native projection format, e.g., equirectangular (ERP), are converted into another projection format, e.g., cubemap (CMP), octahedron (OHP), etc. and frame packed before being fed into existing video codecs. The intermediate projection format is important and would potentially improve the representation efficiency and coding performance. Among all the projection solutions, CMP is very popular and has been widely used in the computer graphics community. The intrinsic rectilinear properties of the CMP format are advantageous for the translational motion model in the modern codec architecture. However, in the CMP representation, the samples on the sphere are not evenly distributed within the faces, resulting in a higher density near the face boundaries and a lower density near the face center. Such non-uniform sampling scheme penalizes the video representation efficiency and degrades the coding performance. Adjusted cubemap projection (ACP) was proposed to address such non-uniform sampling by introducing transform functions to improve the sampling uniformity. However, the transform function parameters in ACP are fixed regardless of the content inside each cube face. In this paper, a generalized hybrid cubemap projection (HCP) is proposed to improve the 360-degree video coding efficiency beyond ACP. HCP is defined by a pair of forward transform and inverse transform functions with a pair of horizontal and vertical transform parameters per cube face. The encoder can choose the optimal sampling for each face by adjusting the parameters in the horizontal and vertical directions based on the 360-degree video content characteristics inside each cube face. In order to maintain the boundary continuities between two neighboring faces, in a 3x2 packing layout, vertical parameter constraints are imposed such that faces in each face-row have the same vertical parameters. The HCP parameters are chosen to minimize the end-to-end weighted conversion error and determined using iterative search between the horizontal and the vertical directions. Significant changes in HCP parameter values can cause drastic change in sampling distribution, and may affect the inter-picture coding efficiency. Therefore, an efficient HCP parameter estimation algorithm is proposed to achieve a better trade-off between the temporal sampling adaptation and the inter-picture prediction efficiency by reducing the temporal variation of HCP parameters. The proposed HCP parameter search algorithm reduces the computational complexity by 5x compared to the exhaustive search method. The HCP parameters are selected by the encoder using the first picture of each Intra Random-Access Point (IRAP) and signalled once per IRAP. In SPS, projection format, frame packing parameters including number of faces in horizontal and vertical directions and each face's position and orientation are signalled. In PPS, the horizontal and vertical HCP parameters in 6-bit precision are encapsulated. The proposed HCP solution is implemented upon JEM-6.0 and 360Lib-3.0 software. Simulation results are reported using the test conditions specified in the JVET Call-for-Evidence (CfE) document. Compared with the CMP and ACP formats, the proposed HCP format demonstrates average 3.0 dB (up to 3.6 dB) and 0.2dB (up to 0.4 dB) End-to-End WS-PSNR improvement for the luma (Y) component, respectively, and average luma (Y) BD-rate reductions of 11.5% (up to 23.0%) and 0.5% (up to 1.0%), respectively. Fanyi Duanmu, Yuwen He, Xiaoyu Xiu, Philippe Hanhart, Yan Ye 0003, Yao Wang 0001 |
DCC | 2 |
| 2018 | 360-Degree Video Quality Evaluationabstract360-degree video is emerging as a new way of offering immersive visual experience. The quality evaluation of 360- degree video is more difficult compared to the quality evaluation of conventional video. However, to ensure successful development of 360-degree video coding technologies, it is essential to precisely measure both objective and subjective quality. In this paper, an overview of the 360-degree video quality evaluation framework established by the joint video exploration team (JVET) of ITU-T VCEG and ISO/IEC MPEG is provided. This framework aims at reproducing the different processes in the 360-degree video processing workflow that are related to coding. The results of different experiments conducted using the JVET framework are reported to illustrate the impact on objective and subjective quality with different projection formats and codecs. Philippe Hanhart, Yuwen He, Yan Ye 0003, Jill M. Boyce, Zhipin Deng, Lidong Xu |
PCS | 2 |
| 2018 | Content-Adaptive 360-Degree Video Coding Using Hybrid Cubemap ProjectionabstractIn this paper, a novel hybrid cubemap projection (HCP) is proposed to improve the 360-degree video coding efficiency. HCP allows adaptive sampling adjustments in the horizontal and vertical directions within each cube face. HCP parameters of each cube face can be adjusted based on the input 360-degree video content characteristics for a better sampling efficiency. The HCP parameters can be updated periodically to adapt to temporal content variation. An efficient HCP parameter estimation algorithm is proposed to reduce the computational complexity of parameter estimation. Experimental results demonstrate that HCP format achieves on average luma (Y) BD-rate reduction of 11.51%, 8.0%, and 0.54% compared to equirectangular projection format, cubemap projection format, and adjusted cubemap projection format, respectively, in terms of end-to-end WS-PSNR. Yuwen He, Xiaoyu Xiu, Philippe Hanhart, Yan Ye 0003, Fanyi Duanmu, Yao Wang 0001 |
PCS | 1 |
| 2018 | Rotational Motion Compensated Prediction in HEVC Based Omnidirectional Video CodingabstractSpherical video is becoming prevalent in virtual and augmented reality applications. With the increased field of view, spherical video needs enormous amounts of data, obviously demanding efficient compression. Existing approaches simply project the spherical content onto a plane to facilitate the use of standard video coders. Earlier work at UCSB was motivated by the realization that existing approaches are suboptimal due to warping introduced by the projection, yielding complex non-linear motion that is not captured by the simple translational motion model employed in standard coders. Moreover, motion vectors in the projected domain do not offer a physically meaningful model. The proposed remedy was to capture the motion directly on the sphere with a rotational motion model, in terms of sphere rotations along geodesics. The rotational motion model preserves the shape and size of objects on the sphere. This paper implements and tests the main ideas from the previous work [1] in the context of a full-fledged, unconstrained coder including, in particular, bi-prediction, multiple reference frames and motion vector refinement. Experimental results provide evidence for considerable gains over HEVC. Bharath Vishwanath, Kenneth Rose, Yuwen He, Yan Ye 0003 |
PCS | 3 |
| 2017 | Geometry Padding for Motion Compensated Prediction in 360 Video Codingabstract360 Video has become popular in recent years, as commercial interests in deploying Virtual Reality (VR) applications rise. This type of video is usually captured using multi-camera arrays, such as the GoPro Omni camera rig. After separate video streams are captured from multiple cameras, image stitching is applied to obtain a spherical representation of the scene, which spans 360 degrees horizontally and 180 degrees vertically, hence the name 360 video. In the existing workflow of 360 spherical video coding, the 360 video is projected onto the 2D plane with a projection format, such as equirectangular (ERP), cubemap (CMP), equal-area (EAP), octahedron (OHP), etc. Most, if not all, of the currently available 360 video content are provided in ERP format defined in longitude and latitude. Projection format conversion may be performed to convert the native ERP format to another format before coding is applied. Some projection formats contain more than one face, for example, CMP projects the sphere onto a cube of six faces or OHP projects the sphere onto an octahedron of eight faces. For these multi-face projection formats, the faces are packed onto a 2D rectangular picture with a frame packing method. For example, the six faces of CMP can be packed with 4x3 configuration, or 3x2 configuration. Finally, the frame packed picture is coded as a 2D conventional video. Existing video codecs are designed only considering conventional 2D video captured on a plane. When motion compensated prediction uses any samples outside of a reference picture's padding will be performed by simply copying the sample values from the picture boundaries. This repetitive padding method is referred as conventional 2D padding method, which is widely used in video coding standards such as H.264, High Efficiency Video Coding (HEVC). However, a 360 video encompasses video information on the whole sphere, and thus intrinsically has a cyclic property. When considering this cyclic property, the reference pictures of a 360 video no longer have boundaries, as the information they contain is all wrapped around a sphere. This cyclic property holds regardless of which projection format or which frame packing is used to represent the 360 video on a 2D plane. The paper presents a new geometry padding method for motion compensated prediction in 360 video coding. Unlike the conventional padding method for 2D video coding, the proposed geometry padding method extends samples outside of a 2D picture's boundaries using neighboring samples on the sphere. The geometry projection format is considered when performing padding. The corresponding sample outside of a face's boundary (which may come from another side in the same face or from another face), is derived with rectilinear projection. Each face is extended with geometry padding separately. When visualized, the extended faces using geometry padding show continuous texture representing natural extension of the texture inside the face. The proposed geometry padding method is implemented in the HEVC reference software HM-16.12 for the ERP and CMP projection formats. In the simulation, a total of sixteen 4K ERP video and eight 8K ERP video are used. For 8K ground truth 8K video, they are converted to 4K video in ERP and CMP projection formats, coded, and converted back to reconstructed 8K video in ERP format. For 4K ground truth video, they are directly coded in 4K ERP, or converted to CMP consisting of 75% of effective samples, coded, and converted back to reconstructed 4K video in ERP format. Then, the end-to-end spherical PSNR (S-PSNR) is calculated between the original 8K or 4K and the reconstructed 8K or 4K ERP video. BD-rate is calculated between the reference unmodified HEVC, which uses the conventional padding method, and HEVC modified with the proposed geometry padding method. Simulation results showed that geometry padding performs better. For 8K sequences, the proposed geometry padding gives on average luma (Y) BD-rate reduction of 0.3% for ERP and 0.8% for CMP, for 4K sequences, the proposed geometry padding gives on average Y BD-rate reduction of 0.2% for ERP and 1.0% for CMP. Comparing the gains in ERP format with the gains in CMP format, the improvement for CMP is larger. This is because CMP has six faces, therefore the improvement from geometry padding method affects more out-of-boundary samples. The proposed geometry padding method is also especially effective for sequences with fast motion. For example, it achieves BD rate reductions of 4.3%, 2.7%, 1.9%, and 2.5% for Glacier, Chairlift, Sb_in_lot, and Driving, respectively. These four sequences are all captured using moving cameras and have fast moving objects. As a result, the sequences contain a lot of across-the-face-boundary motion which can benefit from improved padding method. Detailed simulation results can be found in JVET contribution JVET-D0075 available at http://phenix.int-evry.fr/jvet/doc_end_user/documents/4_Chengdu/wg11/JVET-D0075-v3.zip. Yuwen He, Yan Ye 0003, Philippe Hanhart, Xiaoyu Xiu |
DCC | 1 |
| 2017 | Motion compensated prediction with geometry padding for 360 video codingabstractThe paper presents a new geometry padding method for motion compensated prediction in 360 video coding. Unlike the conventional padding method for 2D video coding, which extends samples outside of a picture's boundary by simply copying (repeating) those at the boundaries, the proposed geometry padding method considers the spherical nature of the 360 video and the specific geometry projection format when extending samples outside of a picture's or a face's boundary. The proposed geometry padding method is implemented in the HEVC reference software HM-16.12 for the equirectangular and cubemap projection formats. Simulation results show that, when compared with the conventional 2D padding method in HEVC, geometry padding performs better; the proposed geometry padding gives on average luma (Y) BD-rate reduction of 0.3% for equirectangular projection format coding and 1.0% for cubemap projection format coding all in terms of the spherical PSNR (S-PSNR) metric. The proposed geometry padding is especially effective for high motion sequences, where up to 4.3% BD-rate reduction can be achieved. Yuwen He, Yan Ye 0003, Philippe Hanhart, Xiaoyu Xiu |
VCIP | 1 |
| 2017 | An evaluation framework for 360-degree video compressionabstract360-degree video is emerging as a new way of offering immersive visual experience. 360-degree video can be viewed on dedicated head mounted devices as well as on conventional 2D displays. Due to increased resolution to support wide field of view, efficient compression of 360-degree video becomes crucial. Whereas it is of significant interest to evaluate how different projection formats impact the compression efficiency of 360-degree video, a main technical challenge is that the input video is captured in a given native projection format, and that the native format has an obvious edge over other projection formats. In this paper, an evaluation framework is proposed to reduce the bias towards the native projection format when comparing different projection formats and their impact on 360-degree video compression. Additionally, a quality metric called area weighted spherical PSNR (AW-SPSNR) is proposed for objective 360-degree video quality evaluation. The proposed evaluation framework was included in the common test conditions defined for the exploration work of 360-degree video compression technologies under the Joint Video Exploration Team (JVET). Xiaoyu Xiu, Yuwen He, Yan Ye 0003, Bharath Vishwanath |
VCIP | 2 |
| 2017 | Overview of Color Gamut ScalabilityabstractDisplays' new rendering capabilities combined with the ever-growing number of video applications have fueled the emergence of new video formats addressing wider color gamut and larger frame size. Thus, the need in scalable compression technology to provide backward compatibility with legacy devices and capitalize on the superior compression performance of High Efficiency Video Coding (HEVC) has increased significantly. This paper gives an overview of the work carried out in the Joint Collaborative Team on Video Coding of ITU-T Study Group 16 (VCEG) and ISO/IEC JTC1/SC29/WG11 motion picture experts group (MPEG) to define scalable extensions of HEVC (SHVC) targeting these market requirements. The color gamut scalability (CGS) tool of SHVC is specially designed to support efficient scalable coding with multiple layers in different color spaces. The genesis of the SHVC-CGS tool is presented and the performance of the various proposals is compared. Finally, the design of the recently adopted SHVC-CGS inter-layer prediction is detailed. The experimental results validate its efficiency in coding video with extended color gamut and high dynamic range. Philippe Bordes, Pierre Andrivon, Xiang Li 0003, Yan Ye 0003, Yuwen He |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2016 | Compression Efficiency Improvement over HEVC Main 10 Profile for HDR and WCG ContentabstractThe paper presents the joint proposal by Arris, Dolby and InterDigital as a response to the Call-for-Evidence of the High Dynamic Range and Wide Color Gamut (HDR/WCG) video compression in MPEG. The joint proposal introduces a set of new HDR coding technologies, including the IPT-PQ color space, the adaptive reshaping process, the color enhancement filters, and the adaptive transfer function. These new coding technologies are applied to the decoded output of an HEVC decoder. Hence, no changes to the lower level logics of the HEVC decoder are required to implement the proposal. Formal subjective tests conducted by MPEG confirmed that the proposal could achieve significant subjective quality improvements over the HEVC Main10 anchors at similar bit rates for HDR/WCG video content. Taoran Lu, Fangjun Pu, Peng Yin 0002, Yuwen He, Louis Kerofsky, Yan Ye 0003, Zhouye Gu, David Baylon |
DCC | 4 |
| 2016 | Recent developments from MPEG in HDR video compressionabstractIn this paper we review the current status and ongoing development of High Dynamic Range and Wide Color Gamut (HDR/WCG) video compression within MPEG. We review how existing MPEG, ITU-R and SMPTE standards may be used for coding HDR content. The history of an exploratory activity within MPEG investigating technologies for improved compression of HDR/WCG content is reviewed. An overview of the MPEG Call for Evidence related to HDR/WCG compression technology is provided. An overview of activities within MPEG related to HDR/WCG coding including progress and a snapshot of ongoing core experiments as of December, 2015 is given. Future outlook for this activity is described. Louis Kerofsky, Yan Ye 0003, Yuwen He |
ICIP | 3 |
| 2016 | Adaptive enhancement filtering for motion compensationabstractThis paper proposes an enhanced motion compensated prediction algorithm for hybrid video coding. The algorithm is built upon the concept of applying adaptive enhancement filtering at the motion compensation stage. For luma, a high-pass filter is applied to the motion compensated prediction signal to recover distorted high-frequency information. For chroma, cross-plane filters are applied to enhance the motion compensated signals by restoring the blurred edges and textures of the chroma planes using the high-frequency information of the luma plane. To verify the effectiveness, the proposed algorithm is implemented on the HM Key Technology Area (HM-KTA) 1.0 platform. Experimental results show that compared to the anchor, the proposed algorithm achieves average Bjentegaard delta (BD) rate savings of 0.4%, 8.8% and 7.4% for Y, Cb and Cr components, respectively. Xiaoyu Xiu, Yuwen He, Yan Ye 0003 |
MMSP | 2 |
| 2016 | Generalized bi-prediction method for future video codingabstractThis paper presents a generalized bi-prediction (GBi) technique to extend the notion of the existing weighted bi-prediction to prediction unit (PU) level. GBi allows bi-prediction weights to be transmitted at the coding unit (CU) level and used for the bi-predicted PUs within that CU. The candidate set of weights in the GBi mode includes a total of 7 weights, including 1/2 used for conventional bi-prediction. An index pointing to the entry of a weight value in the candidate weight set is signaled. At most one index per CU is signaled and the corresponding weight values are shared across all the bi-prediction PUs and all color components in that CU. Besides, GBi can also be extended to inter coding tools, such as, merge mode and affine prediction that support bi-prediction. Experimental results show that GBi performs consistently better than the JEM-2.0 anchor in all rate points, with an average Y Bjentegaard delta rate reduction of 1% and no extra decoding burden introduced. Chun-Chi Chen, Xiaoyu Xiu, Yuwen He, Yan Ye 0003 |
PCS | 3 |
| 2016 | Decoder-side intra mode derivation for block-based video codingabstractThis paper proposes a decoder-side intra mode derivation algorithm for block-based video coding. Instead of explicitly coding intra mode, the algorithm derives intra mode at both encoder and decoder using a template-based method. Based on rate-distortion optimization, the encoder locally determines whether intra mode derivation or intra mode explicit coding is used. Further, as no intra mode is coded, the proposed algorithm is able to more accurately capture the direction of edges in natural videos by increasing the granularity of directional intra predictions with no signaling overhead increase. To verify the effectiveness, the proposed algorithm is implemented on the Joint Exploration Model 2.0 platform. Experimental results show that an average Bjentegaard delta rate saving of 0.9% is achieved by the proposed algorithm. Xiaoyu Xiu, Yuwen He, Yan Ye 0003 |
PCS | 2 |
| 2016 | High Dynamic Range and Wide Color Gamut Video Coding in HEVC: Status and Potential Future EnhancementsabstractAs the video industry begins deployment of ultrahigh-definition TV in both professional and consumer markets, including support for higher dynamic range and wider color gamut services is considered essential within the industry. Higher dynamic range and wider color gamut offer end users a significantly enhanced viewing experience by supporting intensity ranges and colors unattainable in existing distribution ecosystems. In response to this trend, several standardization organizations have launched efforts to better enable these features in both short term and midterm. In this paper, we provide a survey of these standardization activities, with the specific goal of providing a summary of the underlying technologies. Our emphasis is on both existing and potential extensions to the High Efficiency Video Coding standard. Edouard François, Chad Fogg, Yuwen He, Xiang Li 0003, Ajay Luthra, C. Andrew Segall |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Asymmetric 3D Lookup Table Based Color Gamut Scalability in SHVCabstractSHVC is the scalable extension of the latest video coding standard High Efficiency Video Coding (HEVC). Color Gamut Scalability (CGS) refers to a scalable use case in which base layer and enhancement layer have different color gamuts. In this case, special inter-layer prediction is needed to improve coding efficiency in SHVC. In this paper, a solution based on asymmetric 3D lookup table is presented for color gamut scalability. Compared to SHVC without CGS coding tool, the proposed solution provides 9.6% - 16.1% overall luma BD-rate reduction in different test cases. Xiang Li 0003, Jianle Chen, Marta Karczewicz, Yuwen He, Yan Ye 0003, Cheung Auyeung |
DCC | 4 |
| 2015 | Palette-Based Coding in the Screen Content Coding Extension of the HEVC StandardabstractThis paper provides a technical overview of palette-based coding that was adopted into the test model for the screen content coding (SCC) extension of High Efficiency Video Coding (HEVC) standard at the 18th JCT-VC meeting. Key techniques that enable the palette mode to deliver significant coding gains for screen contents are highlighted, including palette table generation, palette table coding, and the coding methods for palette indices and escape colors. Proposed and adopted techniques up to the first version of the working draft of HEVC SCC extension and test model SCM-2.0 are presented. Experimental results are provided to evaluate the performance of the palette mode in the SCC extension of HEVC. Xiaoyu Xiu, Yuwen He, Rajan L. Joshi, Marta Karczewicz, Patrice Onno, Christophe Gisquet, Guillaume Laroche |
DCC | 2 |
| 2014 | Improved Inter-Layer Prediction for the Scalable Extensions of HEVCabstractSummary form only given. Upon the completion of the single-layer H.265/HEVC, scalable extensions of the H.265/HEVC standard, called Scalable High Efficiency Video Coding (SHVC), are currently under development. Compared to the simulcast solution that simply compresses each layer separately, SHVC offers higher coding efficiency by means of inter-layer prediction which is implemented by inserting inter-layer reference (ILR) pictures generated from reconstructed base layer (BL) pictures into the enhancement layer (EL) decoded picture buffer (DPB) for motion-compensated prediction of the collocated pictures in the EL. If the EL has a higher resolution than that of the BL, the reconstructed BL pictures need to be up-sampled to form the ILR pictures. Given that the ILR picture is generated based on the reconstructed BL picture, its suitability for an efficient inter-layer prediction may be limited due to the following reasons. Firstly, quantization is usually applied when coding the BL pictures. Quantization causes the BL reconstructed texture to contain undesired coding artifacts, such as blocking artifacts, ringing artifacts, and color artifacts. Secondly, in case of spatial scalability, a down-sampling process is used to create the BL pictures. To reduce aliasing, the high frequency information in the video signal is typically removed by the down-sampling process. As a result, the texture information in the ILR picture lacks certain high frequency information. In contrast to the ILR picture, the EL temporal reference pictures contain plentiful high frequency information, which could be extracted to enhance the quality of the ILR picture. To further improve the efficiency of inter-layer prediction, a low pass filter may be applied to the ILR picture to alleviate the quantization noise introduced by the BL coding process. In this paper, an ILR enhancement method is proposed to improve the quality of the ILR picture by combining the high frequency information extracted from the EL temporal reference pictures together with the low frequency information extracted from the ILR picture. Experimental results show that the proposed method can significantly increase the ILR efficiency for EL coding, under the Common Test Condition of SHVC, which defines a number of temporal prediction structures called Random Access (RA), Low-delay B (LD-B) and Low-delay P (LD-P), on average the proposed method provides {Y, U, V} BD-rate (BL+EL) gains of {2.0%, 7.1%, 8.2%}, {2.2%, 6.7%, 7.6%} and {4.0%, 7.4%, 8.4%} for RA, LD-B, and LD-P, respectively, in comparison to the performance of the SHVC reference software SHM-2.0. Thorsten Laude, Xiaoyu Xiu, Yuwen He, Yan Ye 0003, Jörn Ostermann |
DCC | 4 |
| 2014 | Scalable extension of HEVC using enhanced inter-layer predictionabstractIn Scalable High Efficiency Video Coding (SHVC), inter-layer prediction efficiency may be degraded because much high frequency information can be removed during: 1) the down-sampling/up-sampling process and, 2) the base layer coding/quantization process. In this paper, we present a method to enhance the quality of the inter-layer reference (ILR) picture by combining the high frequency information from enhancement layer temporal reference pictures with the low frequency information from the up-sampled base layer picture. Experimental results show that on average 3.9% weighted BD-rate gain is achieved compared to SHM-2.0 under SHVC common test conditions. Thorsten Laude, Xiaoyu Xiu, Yuwen He, Yan Ye 0003, Jörn Ostermann |
ICIP | 4 |
| 2014 | Robust 3D LUT estimation method for SHVC color gamut scalabilityabstractColor gamut scalability (CGS) in scalable extensions of High Efficiency Video Coding (SHVC) supports scalable coding with multiple layers in different color spaces. Base layer conveying HDTV video in BT.709 color space and enhancement layer conveying UHDTV video in BT.2020 color space is identified as a practical use case for CGS. Efficient CGS coding can be achieved using a 3D Look-up Table (LUT) based color conversion process. This paper proposes a robust 3D LUT parameter estimation method that estimates the 3D LUT parameters globally using the Least Square method. Problems of matrix sparsity and uneven sample distribution are carefully handled to improve the stability and accuracy of the estimation process. Simulation results confirm that the proposed 3D LUT estimation method can significantly improve coding performance compared with other gamut conversion methods. Yuwen He, Yan Ye 0003 |
VCIP | 1 |
| 2013 | Cross-plane chroma enhancement for SHVC Inter-Layer PredictionabstractThis paper proposes a cross-plane chroma enhancement (CPCE) scheme to enhance the chroma planes of the inter layer reference (ILR) pictures for the Scalable extensions of HEVC (SHVC), the on-going scalable video coding project in JCT-VC. The CPCE scheme restores the blurred edges and textures in the chroma planes using the corresponding information from the luma plane. Experimental results under the SHVC common test conditions show that the average BD-rate reductions for the Cb and Cr chroma planes are as much as -7.5% and -8.5%, respectively, when compared with SHM-1.0. Yan Ye 0003, Yuwen He |
PCS | 3 |
| 2013 | Power aware HEVC streaming for mobileabstractMobile devices, increasingly equipped with high capability processors and connected with fast wireless networks, have become a major consumer of multi-media content. Limited battery life on mobile devices makes power saving a critical factor in delivering a good user experience. This paper proposes a power aware streaming system that combines the emerging High Efficiency Video Coding (HEVC) standard and the Dynamic Adaptive Streaming over HTTP (DASH) standard. The proposed system uses power aware HEVC encoding technologies and client side power adaptation logic to adaptively control power consumption on the client device. The proposed power aware HEVC streaming system can improve quality of experience by setting full-length video playback as client's objective. Demonstration of the proposed power aware HEVC system is available on the ASUS Transformer Xfinity (TF700T) tablet using an ARM processor. Yuwen He, Markus Künstner, Srinivas Gudumasu, Eun-Seok Ryu, Yan Ye 0003, Xiaoyu Xiu |
VCIP | 1 |
| 2010 | Frame Rate Up-Conversion Using Trilateral FilteringabstractFrame rate up-conversion (FRUC) can enhance the visual quality of low frame rate video presented on liquid crystal display. To minimize the difference between a reference block and an interpolated block, an effective FRUC algorithm partitions a large block into several sub-blocks of smaller size and estimates their motions. Motion estimation searches for the block which has the minimum difference (cost) with the processed block in terms of some block matching distortion and motion discontinuity. As convexity and convergence of the cost function are not guaranteed, the computational cost for such motion estimation is usually extensive or unpredictable. In our proposed FRUC method, the two predictions of a frame to be interpolated are generated through shifting its nearest neighbor frames in the previous and following directions with the motion vectors estimated between them. The initial interpolated frame and its pixel's reliability are subsequently estimated from these two predictions. We then apply a trilateral filter on the initial prediction to correct the unreliable pixels and to restore the missing pixels. Our proposed method not only reduces the computation for refining motion vectors, but also suppresses the interpolation noises and misregistration errors. We have conducted extensive experiments and the results show that the proposed algorithm outperforms the existing methods with better objective and subjective visual quality, and achieves about 3 dB on-average peak signal-to-noise ratio improvement. Ci Wang, Lei Zhang 0006, Yuwen He, Yap-Peng Tan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2006 | Block-based Fast Compression for Compound ImagesabstractThis paper presents a novel block-based fast compression (BFC) algorithm for compound images that contain graphics, text and natural images. The images are divided to blocks, which are classified into four different types - smooth blocks, text blocks, hybrid blocks and picture blocks with a fast and effective block-based classification algorithm. Four different coding algorithms are carefully designed for each block type according to their different statistical properties to maximize the compression performance. Simulations show that the BFC algorithm we propose has much lower complexity than DjVu with significant better visual quality at high bit rate, and it also outperforms the popular lossy image coding method JPEG Wenpeng Ding, Dong Liu 0002, Yuwen He, Feng Wu 0001 |
ICME | 3 |
| 2006 | Motion Aligned Spatial Scalable Video CodingabstractA motion aligned spatial scalable video coding scheme (MA-SSC) is proposed in this paper. Different from the traditional spatial scalable coding schemes derived from MPEG-2, in the proposed scheme only one set of intra or inter prediction modes are optimally selected by jointly considering the base and enhancement layers. Thus, it saves one set of macroblock (MB) mode and motion vectors. Moreover, the combined motion estimation can reduce the residual coding bits of the base layer. The MA-SSC and traditional spatial scalable coding schemes are both implemented based on H.264 reference software to evaluate their performance. Simulation results show that the enhancement layer coding efficiency of MA-SSC is up to 0.6dB better than that of the traditional scheme, while the base layer coding efficiency of MA-SSC decreases less than 0.3db compared with the single-layer coding Debing Liu, Yuwen He, Shipeng Li 0001, Debin Zhao, Wen Gao 0001 |
ICME | 2 |
| 2006 | Complexity Scalable 2 : 1 Resolution Downscaling MPEG-2 to WMV Transcoder with Adaptive Error CompensationabstractIn this paper, we focus on 2:1 spatial resolution downscaling transcoding from MPEG-2 to WMV. We propose two architectures (for sequences with or without B-frames respectively) that are unique in their complexity scalability and efficient control over the drifting error, which in return provide a flexible mechanism to achieve desired tradeoff between the complexity and the quality. We achieve resolution downscaling completely in the DCT domain and show that the standard IDCT (as in all the MPEG series standards) can be merged with other DCT-like transform (e.g., the integer transform in WMV) with proper one-time per-element scaling. Extensive experimental results verified the effectiveness of proposed structures against several design objectives such as complexity scalability and performance tradeoffs Guobin Shen, Yuwen He, Wanyong Cao, Shipeng Li 0001 |
ICME | 2 |
| 2006 | Quality-biased rate allocation for compound image coding with block classificationabstractIn this paper, we propose a novel rate allocation method for compound image coding using Quality-biased Rate-Distortion Optimization (QRDO) technique to enhance visual quality. The compound image is divided into 16 times 16 blocks, which are further classified into four types: smooth, text/graphics, continuous tone and hybrid blocks. Four carefully designed coding approaches according to different statistical features are introduced to compress them, and rate-distortion tradeoff with QRDO technique is considered to decide which approach should be used for each block. The quality of "text" regions is specially emphasized during this process by weighting blocks according to their types, which is verified by experimental results Dong Liu 0002, Wenpeng Ding, Yuwen He, Feng Wu 0001 |
ISCAS | 3 |
| 2006 | Complexity scalable MPEG-2 to WMV transcoder with adaptive error compensationabstractIn this paper, we study the problem of video transcoding from MPEG-2 to WMV format, together with several desired functionalities such as bit rate reduction etc. We propose two architectures (for different typical application scenarios) that are unique in their complexity scalability and adaptive drifting error control, which in return provide a mechanism to achieve desired trade-off between the complexity and the quality. A simple model-based rate control algorithm is also presented. We performed extensive experiments for various design targets such as complexity scalability, performance tradeoff, drifting control effect etc. The proposed transcoding architectures can be straightforwardly applied to the MPEG-2 to MPEG-4 transcoding applications due to the significant overlap between the MPEG-4 and WMV coding technology. Guobin Shen, Yuwen He, Wanyong Cao, Shipeng Li 0001 |
ISCAS | 2 |
| 2006 | MPEG-2 to WMV Transcoder With Adaptive Error Compensation and Dynamic SwitchesabstractIn this paper, we study the problem of video transcoding from MPEG-2 to Windows Media Video (WMV) format, together with several desired functionalities such as bit-rate reduction and spatial resolution downscaling. Based on in-depth analysis of error propagation behavior, we propose two architectures (for different typical application scenarios) that are unique in their complexity scalability and adaptive drifting error control, which in return provide a mechanism to achieve a desired tradeoff between complexity and quality. We perform extensive experiments for various design targets such as complexity, scalability, performance tradeoff, and drifting control effect. The proposed transcoding architectures can be straightforwardly applied to the MPEG-2 to MPEG-4 transcoding applications due to the significant overlap between the MPEG-4 and WMV coding technology. Guobin Shen, Yuwen He, Wanyong Cao, Shipeng Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2004 | An unsymmetrical-cross multi-resolution motion search algorithm for MPEG4-AVC/H.264 codingabstractThe new H.264 (MPEG-4 AVC) video-coding standard has a significant performance benefit compared to former standards. Unfortunately, those advanced coding features, including variable block-size motion compensation and multiple reference frames, incur a considerable increase in encoder complexity, especially when using the straightforward full search (FS) algorithm, mainly with regards to motion estimation (ME) and mode decision. We propose a new ME algorithm for fast H.264 coding, utilizing the technique of the generalized motion vector (MV) predictors, named hybrid unsymmetrical-cross multi-resolution grid search (UCMRGS), which takes advantage of the correlation among variable blocks and can considerably reduce the computational cost of ME at the encoder, while at the same time give similar, and, in some cases, better, visual quality compared with the brute force full search algorithm. The proposed algorithms mainly rely upon very robust and reliable predictive techniques with parameters adapted to the local characteristics combined with the UCMRGS pattern Yuwen He, Shiqiang Yang |
ICME | 2 |
| 2003 | Robust video transmission over lossy packet networks using block-based fine granularity scalable coding
Yuwen He, Shiqiang Yang |
VCIP | 1 |
| 2002 | Improved Fine Granular Scalable Coding with Inter-Layer PredictionabstractThis paper proposes an improved fine granular scalable (FGS) coding method with interlayer prediction. There are two important aspects to improving FGS coding efficiency. One is a low bit-rate video coding method and the other is interlayer prediction with enhancement layer reference at base-layer coding. The whole scalable coding performance with the proposed method is greatly enhanced over a wide bandwidth. New spatial and temporal prediction methods are investigated in order to increase FGS base-layer low bit-rate coding efficiency. There are nine kinds of spatial prediction modes for intra coding to exploit the pixels' spatial correlation, including DC and eight directional predictions, and multiple model-based motion prediction is utilized to predict a large irregular motion for inter coding, including a 6-parameter affine model. The reference for motion compensation is adaptively selected from base layer or enhancement layer according to their prediction error. The references are reconstructed at two layers. The references for prediction and reconstruction can be different in order to decrease drifting error due to bit-stream truncation. The coding efficiency of our base layer coding can be comparable with that of latest draft H.26L and holds a compelling improvement compared to MPEG-4. With our proposed scalable coding scheme, the whole FGS coding efficiency can be improved by about 2.0 dB at low bit-rate and 3.0-4.0 dB at medium or high bit-rate. The visual quality of the decoded video is also impressively improved at all decoded bit-rates. Yuwen He, Xuejun Zhao, Yuzhuo Zhong, Shiqiang Yang |
DCC | 1 |
| 2002 | Bilock-based fine granularity scalable video coding for content-aware streamingabstractVideo streaming is becoming more and more popular with widely used hybrid networks. The prime challenge of such applications is to deal with varying transmission bandwidth. This paper proposes a block-based fine granularity scalable (FGS) coding structure, which is a more flexible scalable video coding structure supporting content-aware streaming compared with the MPEG-4 FGS coding structure. The streaming server can conveniently implement content-aware rate allocation or content-based selective enhancement dynamically through user's interaction with the proposed scalable coding structure. Thus the streaming server can have a differentiated delivery strategy according to user's preference. However the uniform rate allocation for bit-stream truncation of the proposed block-based FGS will result in more than 2dB loss by PSNR compared with MPEG-4 FGS within quite a wide range of bit rates. A fast optimal rate allocation method is also proposed to solve this problem in this paper. The coding efficiency is improved, which can be comparable with MPEG-4 FGS coding and is even better (0.5dB) with some sequences at some bit rates. Yuwen He, Shiqiang Yang, Yuzhuo Zhong |
ICIP (2) | 1 |
| 2000 | Region-Based Tracking in Video Sequences Using Planar Perspective Models
Yuwen He, Li Zhao 0006, Shiqiang Yang, Yuzhuo Zhong |
ICMI | 1 |