Ye-Kui Wang

dblp:85/5114 · DBLP profile ↗
← Back
44ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0003-4643-0033ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 38 · 6 first-author · 7 since 2021Systems, architecture and hardware · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2
YearPublicationVenuePosition
2026 An Overview of the JPEG AI Learning-Based Image Coding Standard
abstract
JPEG AI is an emerging learning-based image coding standard developed by Joint Photographic Experts Group (JPEG). The scope of the JPEG AI is the creation of a practical learning-based image coding standard offering a single-stream, compact compressed domain representation, targeting both human visualization and machine consumption. Scheduled for completion in early 2025, the first version of JPEG AI focuses on human vision tasks, demonstrating significant BD-rate reductions compared to existing standards, in terms of MS-SSIM, FSIM, VIF, VMAF, PSNR-HVS, IW-SSIM and NLPD quality metrics. Designed to ensure broad interoperability, JPEG AI incorporates various design features to support deployment across diverse devices and applications. This paper provides an overview of the technical features and characteristics of the JPEG AI standard.
Semih Esenlik, Yaojun Wu 0001, Zhaobin Zhang, Ye-Kui Wang, Kai Zhang 0007, Li Zhang 0006, João Ascenso, Shan Liu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Standardizing Generative Face Video Compression Using Supplemental Enhancement Information
abstract
This paper proposes a Generative Face Video Compression (GFVC) approach using Supplemental Enhancement Information (SEI), where a series of compact spatial and temporal representations of a face video signal (e.g., 2D/3D key-points, facial semantics and compact features) can be coded using SEI messages and inserted into the coded video bitstream. At the time of writing, the proposed GFVC approach using SEI messages has been included into a draft amendment of the Versatile Supplemental Enhancement Information (VSEI) standard by the Joint Video Experts Team (JVET) of ISO/IEC JTC 1/SC 29 and ITU-T SG21, which will be standardized as a new version of ITU-T H.274$|$ISO/IEC 23002-7. To the best of the authors' knowledge, the JVET work on the proposed SEI-based GFVC approach is the first standardization activity for generative video compression. The proposed SEI approach has not only advanced the reconstruction quality of early-day Model-Based Coding (MBC) via the state-of-the-art generative technique, but also established a new SEI definition for future GFVC applications and deployment. Experimental results illustrate that the proposed SEI-based GFVC approach can achieve remarkable rate-distortion performance compared with the latest Versatile Video Coding (VVC) standard, whilst also potentially enabling a wide variety of functionalities including user-specified animation/filtering and metaverse-related applications.
Yan Ye 0003, Jie Chen 0006, Ru-Ling Liao, Shanzhi Yin, Shiqi Wang 0001, Kaifa Yang, Yue Li 0015, Yiling Xu, Ye-Kui Wang, Shiv Gehlot, Guan-Ming Su, Peng Yin 0002, Sean McCarthy, Gary J. Sullivan
IEEE Trans. Multim.10
2025 DPCD: A Quality Assessment Database for Dynamic Point Clouds
abstract
Recently, the advancements in Virtual/Augmented Reality (VR/AR) have driven the demand for Dynamic Point Clouds (DPC). Unlike static point clouds, DPCs are capable of capturing temporal changes within objects or scenes, offering a more accurate simulation of the real world. While significant progress has been made in the quality assessment research of static point cloud, little study has been done on Dynamic Point Cloud Quality Assessment (DPCQA), which hinders the development of quality-oriented applications, such as interframe compression and transmission in practical scenarios. In this paper, we introduce a large-scale DPCQA database, named DPCD, which includes 15 reference DPCs and 525 distorted DPCs from seven types of lossy compression and noise distortion. By rendering these samples to Processed Video Sequences (PVS), a comprehensive subjective experiment is conducted to obtain Mean Opinion Scores (MOS) from 21 viewers for analysis. The characteristic of contents, impact of various distortions, and accuracy of MOSs are presented to validate the heterogeneity and reliability of the proposed database. Furthermore, we evaluate the performance of several objective metrics on DPCD. The experiment results show that DPCQA is more challenge than that of static point cloud. The DPCD, which serves as a catalyst for new research endeavors on DPCQA, is publicly available at https://huggingface.co/datasets/Olivialyt/DPCD.
Qi Yang 0003, Yiling Xu, Zhu Li 0001, Ye-Kui Wang
ICME6
2024 EVAN: Evolutional Video Streaming Adaptation via Neural Representation
abstract
Adaptive bitrate (ABR) using conventional codecs cannot further modify the bitrate once a decision has been made, exhibiting limited adaptation capability. This may result in either overly conservative or overly aggressive bitrate selection, which could cause either inefficient utilization of the network bandwidth or frequent re-buffering, respectively. Neural representation for video (NeRV), which embeds the video content into neural network weights, allows video reconstruction with incomplete models. Specifically, the recovery of one frame can be achieved without relying on the decoding of adjacent frames. NeRV has the potential to provide high video reconstruction quality and, more importantly, pave the way for developing more flexible ABR strategies for video transmission. In this work, a new framework, named Evolutional Video streaming Adaptation via Neural representation (EVAN), which can adaptively transmit NeRV models based on soft actor-critic (SAC) reinforcement learning, is proposed. EVAN is trained with a more exploitative strategy and utilizes progressive playback to avoid re-buffering. Experiments showed that EVAN can outperform existing ABRs with 50% reduction in re-buffering and achieve nearly 20% improvement in users’ quality of experience (QoE).
Mufan Liu, Le Yang 0001, Yiling Xu, Ye-Kui Wang, Jenq-Neng Hwang
ICME4
2021 Extended Dependent Random Access Point Pictures in VVC
abstract
Versatile Video Coding (VVC) is the latest video coding standard, finalized in July 2020. Dependent random access point (DRAP) is supported in HEVC and VVC, through a supplemental enhancement information (SEI) message, for improved coding efficiency in the case of random access. A DRAP picture is an inter-coded picture that can only refer to the most recent intra random access point (IRAP) picture earlier in decoding order. A DRAP picture can be used as a random access point in the bitstream given that the associated IRAP picture is available. Although the DRAP picture increases coding efficiency for random access, the performance is limited since only the associated IRAP picture is allowed to be used as the reference picture for DRAP pictures. In this paper, an extended dependent random access point (EDRAP) picture is proposed to further improve the coding efficiency for random access. Specifically, a particular EDRAP picture only relies on some pictures in a set of pictures consisting of the associated IRAP picture and certain EDRAP pictures between the associated IRAP picture and the particular EDRAP picture in decoding order. EDRAP pictures can be used as random access points as long as the few dependent IRAP or EDRAP pictures are provided. Simulation results demonstrate that the proposed method can achieve 10.2% Bj$\phi$ntegaard-Delta rate saving on average compared to VTM-11.0 under random access configuration. It is noted that the support of EDRAP for VVC, through another SEI message, was adopted by the JVET at its 21stmeeting in January 2021.
Ye-Kui Wang, Li Zhang 0006, Kai Zhang 0007, Zhipin Deng
ICIP2
2021 Developments in International Video Coding Standardization After AVC, With an Overview of Versatile Video Coding (VVC)
abstract
In the last 17 years, since the finalization of the first version of the now-dominant H.264/Moving Picture Experts Group-4 (MPEG-4) Advanced Video Coding (AVC) standard in 2003, two major new generations of video coding standards have been developed. These include the standards known as High Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC). HEVC was finalized in 2013, repeating the ten-year cycle time set by its predecessor and providing about 50% bit-rate reduction over AVC. The cycle was shortened by three years for the VVC project, which was finalized in July 2020, yet again achieving about a 50% bit-rate reduction over its predecessor (HEVC). This article summarizes these developments in video coding standardization after AVC. It especially focuses on providing an overview of the first version of VVC, including comparisons against HEVC. Besides further advances in hybrid video compression, as in previous development cycles, the broad versatility of the application domain that is highlighted in the title of VVC is explained. Included in VVC is the support for a wide range of applications beyond the typical standard- and high-definition camera-captured content codings, including features to support computer-generated/screen content, high dynamic range content, multilayer and multiview coding, and support for immersive media such as 360° video.
Benjamin Bross, Jianle Chen, Jens-Rainer Ohm, Gary J. Sullivan, Ye-Kui Wang
Proc. IEEE5
2021 An Overview of Omnidirectional MediA Format (OMAF)
abstract
During recent years, there have been product launches and research for enabling immersive audio-visual media experiences. For example, a variety of head-mounted displays and 360° cameras are available in the market. To facilitate interoperability between devices and media system components by different vendors, the Moving Picture Experts Group (MPEG) developed the Omnidirectional MediA Format (OMAF), which is arguably the first virtual reality (VR) system standard. OMAF is a storage and streaming format for omnidirectional media, including 360° video and images, spatial audio, and associated timed text. This article provides a comprehensive overview of OMAF.
Miska M. Hannuksela, Ye-Kui Wang
Proc. IEEE2
2021 Overview of the Versatile Video Coding (VVC) Standard and its Applications
abstract
Versatile Video Coding (VVC) was finalized in July 2020 as the most recent international video coding standard. It was developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO/IEC Moving Picture Experts Group (MPEG) to serve an ever-growing need for improved video compression as well as to support a wider variety of today’s media content and emerging applications. This paper provides an overview of the novel technical features for new applications and the core compression technologies for achieving significant bit rate reductions in the neighborhood of 50% over its predecessor for equal video quality, the High Efficiency Video Coding (HEVC) standard, and 75% over the currently most-used format, the Advanced Video Coding (AVC) standard. It is explained how these new features in VVC provide greater versatility for applications. Highlighted applications include video with resolutions beyond standard- and high-definition, video with high dynamic range and wide color gamut, adaptive streaming with resolution changes, computer-generated and screen-captured video, ultralow-delay streaming, 360° immersive video, and multilayer coding e.g., for scalability. Furthermore, early implementations are presented to show that the new VVC standard is implementable and ready for real-world deployment.
Benjamin Bross, Ye-Kui Wang, Yan Ye 0003, Shan Liu 0001, Jianle Chen, Gary J. Sullivan, Jens-Rainer Ohm
IEEE Trans. Circuits Syst. Video Technol.2
2021 The High-Level Syntax of the Versatile Video Coding (VVC) Standard
abstract
Versatile Video Coding (VVC), a.k.a. ITU-T H.266 | ISO/IEC 23090-3, is the new generation video coding standard that has just been finalized by the Joint Video Experts Team (JVET) of ITU-T VCEG and ISO/IEC MPEG at its$19^{\mathrm {th}}$meeting ending on July 1, 2020. This paper gives an overview of the VVC high-level syntax (HLS), which forms its system and transport interface. Comparisons to the HLS designs in High Efficiency Video Coding (HEVC) and Advanced Video Coding (AVC), the previous major video coding standards, are included. When discussing new HLS features introduced into VVC or differences relative to HEVC and AVC, the reasoning behind the design differences and the benefits they bring are described. The HLS of VVC enables newer and more versatile use cases such as video region extraction, composition and merging of content from multiple coded video bitstreams, and viewport-adaptive 360° immersive media.
Ye-Kui Wang, Robert Skupin, Miska M. Hannuksela, Sachin Deshpande, Hendry, Virginie Drugeon, Rickard Sjöberg, Byeongdoo Choi, Vadim Seregin, Yago Sánchez de la Fuente, Jill M. Boyce, Wade Wan, Gary J. Sullivan
IEEE Trans. Circuits Syst. Video Technol.1
2020 Video Codec Using Flexible Block Partitioning and Advanced Prediction, Transform and Loop Filtering Technologies
abstract
This paper describes a joint response to the Call for Proposals by Samsung, Huawei, GoPro, and HiSilicon on Video Compression with Capability beyond HEVC/H.265, jointly issued by ITU-T SG16 Q.6 (VCEG) and ISO/IEC JTC1/SC29/WG11 (MPEG). In the proposed codec, the coding framework supports hierarchical splitting with binary and ternary trees and flexible coding order representations. Additionally, novel compression tools on inter/intra prediction, in-loop filtering, and entropy coding have been proposed. The proposed compression scheme provides significantly higher compression capability than the state-of-the-art HEVC/H.265 standard for SDR (Standard Dynamic Range) category while maintaining complexity acceptable for emerging applications. When all the proposed algorithmic tools are used, the proposed video codec achieves approximately 40% bit-saving for the SDR cetegory on average compared to HEVC/H.265 anchor.
Kiho Choi, Jianle Chen, Haitao Yang 0001, Woongil Choi, Sergey Ikonin, Yinji Piao, Semih Esenlik, Minsoo Park, Ye-Kui Wang, Narae Choi, Yin Zhao, Seungsoo Jeong, Anish Tamse, Alexey Filippov, Heechul Yang, Junghye Min, Roman Chernyak, Bora Jin, Anand Meher Kotra, Sunil Lee, Han Gao 0001, Chanyul Kim, Timofey Solovyev, Kwangpyo Choi, Vasily Rufitskiy, Maxim Sychev, Jeonghoon Park
IEEE Trans. Circuits Syst. Video Technol.10
2019 An Overview of the OMAF Standard for 360° Video
abstract
Omnidirectional MediA Format (OMAF) is arguably the first virtual reality (VR) system standard, recently developed by the Moving Picture Experts Group (MPEG). OMAF defines a media format that enables omnidirectional media applications, focusing on 360° video, images, and audio, as well as the associated timed text, supporting three degrees of freedom (3DOF). This paper gives an overview of the first edition of the OMAF standard.
Miska M. Hannuksela, Ye-Kui Wang, Ari Hourunranta
DCC2
2017 Spatially Scalable HEVC for Layered Division Multiplexing in Broadcast
abstract
Recent broadcast standards support Layered Division Multiplexing (LDM) to achieve graceful degradation as signal quality degrades at the receiver. LDM is accomplished by using different constellations within the same Radio Frequency (RF) spectrum. LDM thus enables delivering multiple service tiers in a single broadcast channel. LDM when used in conjunction with scalable source coding codecs such as the Scalable extension of High Efficiency Video Coding (SHVC), further helps improve overall spectrum utilization and efficiency. In this paper we investigate a 2-tier broadcast LDM based service with one service tier aimed at lower video resolution such as 540p, 720p, 1080p for a mobile receiver (smaller/indoor antenna) and the other service tier targeting twice the video resolution of the lower tier, for stationary receivers (larger/outdoor antenna). The primary contribution of this paper is to identify 2-tier transmission configurations of interest to broadcasters and compare the spectrum and bitrate coding efficiency gains of an SHVC-based multi-tier service versus a simulcast (single layer) based multi-tier service for an Advanced Television Systems Committee (ATSC) 3.0 transmission system. Bitrate savings ranging from 38% and 57% is observed for the SHVC based layered system. For large coverage and pedestrian with a receiver test scenarios, channel utilization savings ranging from 23% to 46% is observed. For mobile and tablet in bedroom scenarios a smaller broadcast bandwidth savings ranging from 6% to 9% is observed.
Kiran M. Misra, C. Andrew Segall, Jie Zhao 0007, Seung-Hwan Kim 0001, Joan Llach, Alan Stein, John Stewart, Hendry, Ye-Kui Wang, Yan Ye 0003
DCC9
2017 Standardization status of 360 degree video coding and delivery
abstract
The emergence of consumer level capturing and display devices for 360 degree video creates new and promising segments in entertainment, education, professional training, and other markets. In order to avoid market fragmentation and ensure interoperability of 360 degree video ecosystems, industry and academia cooperate in standardization efforts in this field. In the video coding domain, 360 degree video invalidates many established procedures, e.g., concerning evaluation of the visual quality, while the specific content characteristics offer potential for higher compression efficiency beyond the current standards. Likewise, 360 degree video puts stricter demands on the system level aspects of transmission but may also offer the potential to enhance existing transport schemes. The Joint Collaborative Team on Video Coding (JCT-VC) as well as the Joint Video Exploration Team (JVET) already started investigations into 360 degree video coding while numerous activities in the Systems subgroup of the Moving Picture Experts Group (MPEG) started to investigate application requirements and delivery aspects of 360 degree video. This paper reports on the current status of the outlined standardization efforts.
Robert Skupin, Yago Sánchez de la Fuente, Ye-Kui Wang, Miska M. Hannuksela, Jill M. Boyce, Mathias Wien
VCIP3
2016 Overview of the Multiview and 3D Extensions of High Efficiency Video Coding
abstract
The High Efficiency Video Coding (HEVC) standard has recently been extended to support efficient representation of multiview video and depth-based 3D video formats. The multiview extension, MV-HEVC, allows efficient coding of multiple camera views and associated auxiliary pictures, and can be implemented by reusing single-layer decoders without changing the block-level processing modules since block-level syntax and decoding processes remain unchanged. Bit rate savings compared with HEVC simulcast are achieved by enabling the use of inter-view references in motion-compensated prediction. The more advanced 3D video extension, 3D-HEVC, targets a coded representation consisting of multiple views and associated depth maps, as required for generating additional intermediate views in advanced 3D displays. Additional bit rate reduction compared with MV-HEVC is achieved by specifying new block-level video coding tools, which explicitly exploit statistical dependencies between video texture and depth and specifically adapt to the properties of depth maps. The technical concepts and features of both extensions are presented in this paper.
Gerhard Tech, Ying Chen 0011, Karsten Müller 0001, Jens-Rainer Ohm, Anthony Vetro, Ye-Kui Wang
IEEE Trans. Circuits Syst. Video Technol.6
2014 Motion Hooks for the Multiview Extension of HEVC
abstract
MV-HEVC refers to the multiview extension of High Efficiency Video Coding (HEVC). At the time of writing, MV-HEVC was being developed by the Joint Collaborative Team on 3D Video Coding Extension Development (JCT-3V) of International Organization for Standardization (ISO)/International Electrotechnical Commission (IEC) Moving Picture Experts Group and ITU-T VCEG. Before HEVC itself was technically finalized in January 2013, the development of MV-HEVC had already started and it was decided that MV-HEVC would only contain high-level syntax changes compared with HEVC, i.e., no changes to block-level processes, to enable the reuse of the first-generation HEVC decoder hardware as is for constructing an MV-HEVC decoder with only firmware changes corresponding to the high-level syntax part of the codec. Consequently, any block-level process that is not necessary for HEVC itself but on the other hand is useful for MV-HEVC can only be enabled through so-called hooks. Motion hooks refer to techniques that do not have a significant impact on the HEVC single-view version 1 codec and can mainly improve MV-HEVC. This paper presents techniques for efficient MV-HEVC coding by introducing hooks into the HEVC design to accommodate inter-view prediction in MV-HEVC. These hooks relate to motion prediction, hence named motion hooks. Some of the motion hooks developed by the authors have been adopted into HEVC during its finalization. Simulation results show that the proposed motion hooks provide on average 4% of bitrate reduction for the views coded with inter-view prediction.
Ying Chen 0011, Li Zhang 0006, Vadim Seregin, Ye-Kui Wang
IEEE Trans. Circuits Syst. Video Technol.4
2012 System Layer Integration of High Efficiency Video Coding
abstract
This paper describes the integration of High Efficiency Video Coding (HEVC) into end-to-end multimedia systems, formats, and protocols such as Real-time transport Protocol, the transport stream of the MPEG-2 standard suite, and dynamic adaptive streaming over the Hypertext Transport Protocol. This paper gives a brief overview of the high-level syntax of HEVC and the relation to the Advanced Video Coding standard (H.264/AVC). A section on HEVC error resilience concludes the HEVC overview. Furthermore, this paper describes applications of video transport and delivery such as broadcast, television over the Internet Protocol, Internet streaming, video conversation, and storage as provided by the different system layers.
Thomas Schierl, Miska M. Hannuksela, Ye-Kui Wang, Stephan Wenger
IEEE Trans. Circuits Syst. Video Technol.3
2012 Overview of HEVC High-Level Syntax and Reference Picture Management
abstract
The increasing proportion of video traffic in telecommunication networks puts an emphasis on efficient video compression technology. High Efficiency Video Coding (HEVC) is the forthcoming video coding standard that provides substantial bit rate reductions compared to its predecessors. In the HEVC standardization process, technologies such as picture partitioning, reference picture management, and parameter sets are categorized as “high-level syntax.” The design of the high-level syntax impacts the interface to systems and error resilience, and provides new functionalities. This paper presents an overview of the HEVC high-level syntax, including network abstraction layer unit headers, parameter sets, picture partitioning schemes, reference picture management, and supplemental enhancement information messages.
Rickard Sjöberg, Ying Chen 0011, Akira Fujibayashi, Miska M. Hannuksela, Jonatan Samuelsson, Thiow Keng Tan, Ye-Kui Wang, Stephan Wenger
IEEE Trans. Circuits Syst. Video Technol.7
2010 Multiple Description Video Coding With H.264/AVC Redundant Pictures
abstract
Multiple description coding offers interesting solutions for error resilient multimedia communications as well as for distributed streaming applications. In this letter, we propose a scheme based on H.264/AVC for encoding of image sequences into multiple descriptions. The pictures are split into multiple coding threads. Redundant pictures are inserted periodically in order to increase the resilience to loss and to reduce the error propagation. They are produced with different reference frames than the corresponding primary pictures. We show, given the channel conditions, how to optimally allocate the rates to primary and redundant pictures, such that the total distortion at the receiver is minimized. Extensive experiments demonstrate that the proposed scheme outperforms baseline solutions based on loss and content-adaptive intra coding. Finally, we show how to further reduce the distortion by efficient combination of primary and redundant pictures, if both are available at the decoder.
Ivana Radulovic, Pascal Frossard, Ye-Kui Wang, Miska M. Hannuksela, Antti Hallapuro
IEEE Trans. Circuits Syst. Video Technol.3
2009 Spatial transcoding from Scalable Video Coding to H.264/AVC
abstract
Scalable Video Coding (SVC) is backwards compatible to H.264/AVC in the sense that the base layer sub-bitstream is decodable by an H.264/AVC decoder. However, there are applications wherein it is desirable for an H.264/AVC decoder to obtain a higher resolution video representation than the base layer within SVC. In order to fulfill the needs of such application scenarios, transcoding of SVC enhancement layers to H.264/AVC is required. This paper presents a transcoding scheme that is capable of transcoding a spatial scalable SVC bitstreams to H.264/AVC bitstreams that provide high resolution than the H.264/AVC compliant base layer. To reduce the complexity at the transcoder, a fast mode decision (MD) process is proposed, wherein the original SVC macroblock coding modes and motion information are reused as much as possible. Experimental results show that proposed scheme performs elegantly compared with full-decoding-and-encoding transcoding with low computational complexity.
Ye-Kui Wang, Ying Chen 0011, Houqiang Li
ICME2
2009 Regionally Adaptive Filtering for Asymmetric Stereoscopic Video Coding
abstract
In asymmetric stereoscopic video coding, one view can be coded in a lower resolution of the other. In this scenario, stereoscopic video can be compressed with only moderately increased bandwidth and complexity compared to 2D monoview video coding. The subjective quality degradation of this scenario can be negligible compared to coding two views with original resolution. The low-resolution view can be predicted from the high-resolution view to achieve higher coding efficiency. In this paper, a regionally adaptive filtering algorithm is proposed to generate a predictor of a macroblock (MB) or MB partition of the low-resolution view from the high-resolution view. Different filters are applied for different picture regions. Disparity motion matching and clustering are applied in the encoder for generation of regionally adaptive filters. Simulation results show that the proposed algorithm results in up to 27% bit-rate saving compared with methods without adaptive filtering.
Ying Chen 0011, Ye-Kui Wang, Moncef Gabbouj, Miska M. Hannuksela
ISCAS2
2009 Joint Texture and Depth Map Video Coding based on the Scalable Extension of H.264/AVC
abstract
Depth-Image-Based Rendering (DIBR) is widely used for view synthesis in 3D video applications. Compared with traditional 2D video applications, both the texture video and its associated depth map are required for transmission in a communication system that supports DIBR. To efficiently utilize limited bandwidth, coding algorithms, e.g. the Advanced Video Coding (H.264/AVC) standard, can be adopted to compress the depth map using the 4:0:0 chroma sampling format. However, when the correlation between texture video and depth map is exploited, the compression efficiency may be improved compared with encoding them independently using H.264/AVC. A new encoder algorithm which employs Scalable Video Coding (SVC), the scalable extension of H.264/AVC, to compress the texture video and its associated depth map is proposed in this paper. Experimental results show that the proposed algorithm can provide up to 0.97 dB gain for the coded depth maps, compared with the simulcast scheme, wherein texture video and depth map are coded independently by H.264/AVC.
Siping Tao, Ying Chen 0011, Miska M. Hannuksela, Ye-Kui Wang, Moncef Gabbouj, Houqiang Li
ISCAS4
2009 Efficient hierarchical inter picture coding for H.264/AVC baseline profile
abstract
Bi-predictive (B) slices are not supported in the Baseline profile of the Advanced Video Coding (H.264/AVC) standard, which results in a decreased coding efficiency compared with other profiles supporting B slices. However, many application standards, such as the mobile multimedia services specified by the Third Generation Partnership Project (3GPP), use only the Baseline profile for H.264/AVC. Therefore, it is worth investigating H.264/AVC coding when only intra (I) and inter (P) slices are supported. In this paper, a content-adaptive Quantization Parameter (QP) cascading scheme for the hierarchical P coding method compatible with Baseline profile of H.264/AVC is proposed. The proposed method is based on a picture-level QP optimization. The proposed method has a significantly better rate-distortion performance than the traditional IPPP coding structure and outperforms hierarchical P coding methods using fixed delta QP settings between temporal levels noticeably with up to 0.53 dB gain in average luminance Peak Signal-to-Noise Ratio (PSNR).
Weixing Wan, Ying Chen 0011, Ye-Kui Wang, Miska M. Hannuksela, Houqiang Li, Moncef Gabbouj
PCS3
2009 Error Resilient Coding and Error Concealment in Scalable Video Coding
abstract
Scalable video coding (SVC), which is the scalable extension of the H.264/AVC standard, was developed by the Joint Video Team (JVT) of ISO/IEC MPEG (Moving Picture Experts Group) and ITU-T VCEG (Video Coding Experts Group). SVC is designed to provide adaptation capability for heterogeneous network structures and different receiving devices with the help of temporal, spatial, and quality scalabilities. It is challenging to achieve graceful quality degradation in an error-prone environment, since channel errors can drastically deteriorate the quality of the video. Error resilient coding and error concealment techniques have been introduced into SVC to reduce the quality degradation impact of transmission errors. Some of the techniques are inherited from or applicable also to H.264/AVC, while some of them take advantage of the SVC coding structure and coding tools. In this paper, the error resilient coding and error concealment tools in SVC are first reviewed. Then, several important tools such as loss-aware rate-distortion optimized macroblock mode decision algorithm and error concealment methods in SVC are discussed and experimental results are provided to show the benefits from them. The results demonstrate that PSNR gains can be achieved for the conventional inter prediction (IPPP) coding structure or the hierarchical bi-predictive (B) picture coding structure with large group of pictures size, for all the tested sequences and under various combinations of packet loss rates, compared with the basic joint scalable video model (JSVM) design applying no error resilient tools at the encoder and only picture copy error concealment method at the decoder.
Ying Chen 0011, Ye-Kui Wang, Houqiang Li, Miska M. Hannuksela, Moncef Gabbouj
IEEE Trans. Circuits Syst. Video Technol.3
2009 Error Resilient Video Coding Using Redundant Pictures
abstract
This paper presents several error resilient video coding methods based on redundant pictures. We combine redundant picture coding with reference picture selection and reference picture list reordering to prevent error propagation in motion compensated video coding. A hierarchical redundant picture allocation method is employed to make a tradeoff between error resilience and coding efficiency. For improved end-to-end rate-distortion performance in packet loss environment, three adaptive redundant picture allocation methods are further developed, utilizing characteristics of the input video content. Simulation results show that the adaptive redundant picture coding methods can achieve average luma peak-signal-to-noise improvements up to 3.5 dB compared to the loss-aware rate distortion optimized intra macroblock refresh algorithm implemented in the H.264/AVC Joint Model (JM). The proposed redundant picture coding methods are standard-compliant and do not introduce any additional end-to-end delay, therefore suit for low-delay applications such as video telephony and video conferencing. Due to the good error resilience performance, some of the proposed redundant picture coding methods have been adopted and integrated into the JM.
Chunbo Zhu, Ye-Kui Wang, Miska M. Hannuksela, Houqiang Li
IEEE Trans. Circuits Syst. Video Technol.2
2008 Picture-level adaptive filter for asymmetric stereoscopic video
abstract
In asymmetric stereoscopic video coding, one view is coded in a quarter of the resolution of the other and the low- resolution view is predicted from the high-resolution view. This way, stereoscopic video effect could be achieved with only moderately increased bandwidth and complexity. Inter-view prediction tools for generating the predictor of a maroblock (MB) or MB partition in the low-resolution view from the high-resolution view play a vital role for coding efficiency in asymmetric video coding. In this paper, we propose a method that applies an adaptive filter to generate picture-level adaptive inter-view predictors for MBs or MB partitions. At the encoder, a low complexity preprocessing module is built to find out the filters. Simulation results show that the proposed method provides a bit-rate saving of 26% at maximum and 5% on average.
Ying Chen 0011, Ye-Kui Wang, Miska M. Hannuksela, Moncef Gabbouj
ICIP2
2008 Low-complexity asymmetric multiview video coding
abstract
Multiview video coding (MVC) is currently under development by the Joint Video Team (JVT) as an extension to Advanced Video Coding (H264/AVC). Based on the suppression theory in binocular vision, the fidelity of one of the two views of a stereoscopic display can be reduced without noticeable degradation of subjective quality. Thus, in MVC, a subset of views can be coded with lower spatial resolution at negligible cost to subjective quality. Due to different resolutions, a downsampling process is required in an MVC decoder in order to enable motion compensation (MC) between views. In this paper, a low-complexity MC algorithm is proposed for MVC to enable inter-view prediction between pictures with different resolutions. It requires lower memory consumption and lower computational complexity compared with the conventional downsampled inter-view prediction, while providing comparable efficiency, as shown by the simulation results.
Ying Chen 0011, Shujie Liu 0001, Ye-Kui Wang, Miska M. Hannuksela, Houqiang Li, Moncef Gabbouj
ICME3
2008 Single-loop decoding for multiview video coding
abstract
Multiview video coding (MVC) is currently being standardized by the Joint Video Team as an extension of H264/AVC. When an MVC bitstream is decoded, some views (named target views) are to be displayed; some other views (named dependent views) may not be displayed but are needed for inter-view prediction of the target views. The original MVC design requires pictures of the dependent views to be fully decoded and stored. This entails both high decoding complexity and high memory consumption for the pictures in the views which are not intended for display, particularly when the number of dependent views is large. In this paper, a single-loop decoding (SLD) scheme is introduced to address these disadvantages. SLD requires only partial decoding of pictures in dependent views and thus significantly reduces decoding complexity and memory consumption. The proposed method is based on the so-called motion skip, wherein inter-view motion and coding mode prediction is exploited. Experimental results show that compared to coding schemes that require comparable complexity, significant compression gain can be achieved. For example, 25% bit-rate saving on average can be obtained compared to simulcast. Simulation results also show that the proposed SLD scheme provides a substantial reduction of complexity and memory size, at the expense of only a minor compression efficiency loss, compared with multiple-loop decoding MVC schemes.
Ying Chen 0011, Ye-Kui Wang, Miska M. Hannuksela, Moncef Gabbouj
ICME2
2008 Priority-based template matching intra prediction
abstract
Template matching intra prediction has been reported earlier, to generate intra prediction signal from wider reconstructed areas compared with H.264/AVC intra prediction. In this paper an improved intra prediction scheme which introduces priority-guided template matching is presented. The proposed method first calculates a priority value of each pixel on the border between the current to-be-predicted block and the reconstructed or predicted area, and then performs template matching for the template centered on the pixel with the highest priority among all the border pixels, in an iterative manner. Simulation results demonstrate that the proposed algorithm can save about 2% bitrate than the non-priority algorithm.
Ye-Kui Wang, Houqiang Li
ICME2
2008 Frame loss error concealment for multiview video coding
abstract
The Multiview Video Coding (MVC) standard is currently under development by the Joint Video Team as an extension of the Advanced Video Coding (H.264/AVC) standard. An MVC encoder compresses more than one viewpoint of a scene captured by different cameras. Redundancies between views can be used for inter-view prediction in encoding as well as error concealment in decoding. In this paper, a new algorithm utilizing motion information of pictures from other views to conceal a lost picture is proposed. The algorithm first derives motion information for a lost picture based on motion fields of pictures in adjacent views. Then, traditional motion compensation is invoked within the view containing the lost picture to derive a concealed frame. Experimental results show that the proposed algorithm can improve video quality with a negligible computational complexity overhead compared to simple temporal error concealment algorithms.
Shujie Liu 0001, Ying Chen 0011, Ye-Kui Wang, Moncef Gabbouj, Miska M. Hannuksela, Houqiang Li
ISCAS3
2008 Evaluation of Error Resilience Mechanisms for 3G Conversational Video
abstract
Communication in 3G networks may experience packet losses due to transmission errors on the wireless link(s) which may severely impact the quality of video services, with conversational video being most challenging to repair due to tighter delay constraints. Many error resilience mechanisms have been developed that can be applied at the source (codec) level and transport/application layer to address these challenges. Their respective performance varies depending on the network conditions. This paper analyzes and compares the performance of four error resilience mechanisms under different realistic wireless link conditions: selective retransmissions, slice size adaptation, reference picture selection, and unequal error protection using packet-based forward error correction. We derive suggestions for the applicability of the individual mechanisms.
Jegadish Devadoss, Varun Singh, Jörg Ott, Ye-Kui Wang, Igor D. D. Curcio
ISM5
2008 Error resilient transcoding of Scalable Video bitstreams
abstract
We propose in this paper a novel error resilient transcoding scheme that can be placed at the boundary between wired and wireless networks via heterogeneous network links. This error resilient transcoder shall seamlessly complement the standard Scalable Video Coding (SVC) bitstream to offer additional error resilient adaptation capability for receiving devices. The novel error resilient transcoding scheme consists of three different modules; each is designed to meet various levels of complexity need. The three modules are all based on the Loss-Aware Rate-Distortion Optimization (LA-RDO) mode decision algorithm we have previously developed for SVC. However, each individual module can be tailored to different complexity requirements depending on whether and how the LA-RDO mode decision is implemented. Another innovation of this approach is the design of a fast rate control algorithm in order to maintain consistent bitrates between input and output of the transcoder. This rate control algorithm only needs picture-level bit information for training target quantization parameters. Simulation results demonstrate that, comparing with standard SVC, the proposed approach is able to achieve up to 4 dB gain for the enhancement layer video and up to 1 dB gain for the base layer video.
Houqiang Li, Ye-Kui Wang, Chang Wen Chen
MMSP3
2007 Adaptive Redundant Picture for Error Resilient Video Coding
abstract
We present several efficient adaptive redundant picture coding methods for error resilient video coding. In our previous work, redundant picture coding in combining with reference picture selection, reference picture list reordering and hierarchical redundant picture allocation was proposed. This paper investigates how to allocate redundant pictures more efficiently according to the content characteristics of the primary pictures. Simulation results show that the adaptive redundant picture coding methods can achieve average PSNR improvements around 2 to 4 dB compared to the loss-aware rate distortion optimized (LA-RDO) intra macroblock refresh implemented in H.264/AVC Joint Model (JM). This paper also asserts that the methods do not introduce any additional end-to-end delay, therefore suit for low-delay applications such as video telephony and video conferencing, which demand better error resilience than other applications.
Chunbo Zhu, Ye-Kui Wang, Houqiang Li
ICIP (4)2
2007 System and Transport Interface of SVC
abstract
Scalable video coding (SVC) and transmission has been a research topic for many years. Among other objectives, it aims to support different receiving devices, perhaps connected through a heterogeneous network structure, using a single bit stream. Earlier attempts of standardized scalable video coding, for example in MPEG-2, H.263, or MPEG-4 Visual, have not been commercially successful. Nevertheless, the Joint Video Team has recently focused on the development of the scalable video extensions of H.264/AVC, known as SVC. Some of the key problems of older scalable compression techniques have been solved in SVC and, at the same time, new and compelling use cases for SVC have been identified. While it is certainly important to develop coding tools targeted at high coding efficiency, the design of the features of the interface between the core coding technologies and the system and transport are also of vital importance for the success of SVC. Only through this interface, and novel mechanisms defined therein, applications can take advantage of the scalability features of the coded video signal. This paper provides an overview of the system interface features defined in the SVC specification. We discuss, amongst other features, bit stream structure, extended network abstraction layer (NAL) unit header, and supplemental enhancement information (SEI) messages related to scalability information.
Ye-Kui Wang, Miska M. Hannuksela, Stéphane Pateux, Alexandros Eleftheriadis, Stephan Wenger
IEEE Trans. Circuits Syst. Video Technol.1
2007 Transport and Signaling of SVC in IP Networks
abstract
The transport of scalable media, and in particular of scalable video conforming to the forthcoming Scalable Video Coding (SVC) technology, presents challenges not only in the video compression technology, but also in transport and signaling. This paper discusses the current status of standardization of the support for scalable media, and SVC in particular, over IP based networks. Both the transport of SVC over the Real-time Transport Protocol (RTP), and the signaling support-namely the additional mechanisms in the Session Description Protocol (SDP)-are covered. As it turns out, the support of SVC over RTP is not quite as straightforward as that of nonscalable video bit streams. Specifically, the signaling architecture requires an almost complete overhaul, and new protocol mechanisms need to be introduced into the packetization.
Stephan Wenger, Ye-Kui Wang, Thomas Schierl
IEEE Trans. Circuits Syst. Video Technol.2
2006 Error Resilient Mode Decision in Scalable Video Coding
abstract
Error resilient macroblock mode decision has been extensively investigated in the literature for single-layer video coding, for which error resilient mode decision is also called as intra refresh. In this paper, we present a loss-aware rate-distortion optimized macroblock mode decision algorithm for scalable video coding, wherein more macroblock coding modes than intra and inter are involved. Thanks to the good performance, the proposed method has been adopted into the joint scalable video model by the joint video team.
Ye-Kui Wang, Houqiang Li
ICIP2
2006 System and Transport Interface of H.264/AVC Scalable Extension
abstract
The scalable extension of H.264/AVC, known as scalable video coding or SVC, has recently been the main focus of the Joint Video Team. The higher level syntax of SVC follows the design principles of H.264/AVC. This allows the work towards an optimized file format and RTF payload format to be conducted in parallel with the core SVC specification. This paper provides an overview of the system and transport interface design of SVC, including the SVC high-level syntax and functionalities, SVC file format, and SVC RTF payload format.
Ye-Kui Wang, Stephan Wenger, Miska M. Hannuksela
ICIP1
2006 Error Resilient Video Coding using Redundant Pictures
abstract
Coding of redundant pictures is supported in the latest international video coding standard H.264 (also known as MPEG-4 part 10 or AVC). This paper proposes a standard-compliant way to encode and decode redundant pictures for improved error resilience. The method is based on a combination of picture-level reference picture selection, reference picture list ordering, and a hierarchical allocation of redundant pictures, which can efficiently prevent temporal error propagation without relying on feedback information. Simulation results show that the method outperforms the optimal loss-aware rate-distortion optimized intra refresh method. The proposed algorithm has been adopted into the H.264 joint model.
Chunbo Zhu, Ye-Kui Wang, Miska M. Hannuksela, Houqiang Li
ICIP2
2006 AVS-M: From Standards to Applications
Ye-Kui Wang
J. Comput. Sci. Technol.1
2004 Isolated regions in video coding
abstract
Different types of prediction are applied in modern video coding. While predictive coding improves compression efficiency, the propagation of transmission errors becomes more likely. In addition, predictive coding brings difficulties to other aspects of video coding, including random access, parallel processing, and scalability. In order to combat the negative effects, video coding schemes introduce mechanisms such as slices and intracoding, to limit and break the prediction. This paper proposes the use of the isolated regions coding tool that jointly limits in-picture prediction and interprediction on a region-of-interest basis. The tool can be used to provide random access points from non-intrapictures and to respond to intrapicture update requests. Furthermore, it can be applied as an error-robust macroblock mode decision method and can be used in combination with unequal error protection. Finally, it enables mixing of scenes, which is useful in coding of masked scene transitions.
Miska M. Hannuksela, Ye-Kui Wang, Moncef Gabbouj
IEEE Trans. Multim.2
2003 Random access using isolated regions
abstract
Random access is a desirable feature in many video communication systems. Intra pictures is conventionally used as random access points, but correct picture content is recovered gradually within a range of pictures starting from a non-intra random access point. This paper proposes the use of the isolated regions technique for gradual decoder refresh and presents how the proposed method can be used in the upcoming ITU-T recommendation H.264, also known as MPEG-4 part 10 or advanced video coding. The presented simulations reveal that the proposed method outperforms intra-picture-based random access points in error-prone network conditions. It is also shown that the proposed method is more flexible and suits packet-based transmission better compared to progressively located intra-coded slices.
Miska M. Hannuksela, Ye-Kui Wang, Moncef Gabbouj
ICIP (3)2
2002 Coding of faded scene transitions
abstract
Coding of a scene transition is often a challenging problem, from the compression efficiency point of view, because motion compensation may not be a powerful enough method to represent changes between pictures in the transition. This paper proposes a overlay coding technique for coding faded scene transitions. As shown by extensive simulations, over 50% bit-rate savings in both cross-fades and through-black fades compared to earlier techniques can be achieved. Overlay coding suits situations where video is edited manually or automatically.
Dong Tian, Miska M. Hannuksela, Ye-Kui Wang, Moncef Gabbouj
ICIP (2)3
2002 Sub-picture: ROI coding and unequal error protection
abstract
Region-of-interest coding and unequal error protection are two important tools in video communication systems to improve the received visual quality. One common property of the two techniques is that unequal coding or transmission is applied to improve the quality of the most important parts of images. The proposed sub-picture coding technique facilitates both region-of-interest coding and unequal error protection by partitioning images to regions of interest and separating the corresponding coded data units from each other. Simulation results show that the overall subjective quality is considerably improved compared to the conventional coding schemes.
Ye-Kui Wang, Miska M. Hannuksela, Moncef Gabbouj
ICIP (3)1
2002 The error concealment feature in the H.26L test model
abstract
This paper presents the error concealment (EC) feature implemented by the authors in the test model of the draft ITU-T video coding standard H.26L. The selected EC algorithms are based on weighted pixel value averaging for INTRA. pictures and boundary-matching-based motion vector recovery for INTER pictures. The specific concealment strategy and some special methods, including handling of B-pictures, multiple reference frames and entire frame losses, are described. Both subjective and objective results are given based on simulations under Internet conditions. The feature was adopted and is now included in the latest H.26L reference software TML-9.0.
Ye-Kui Wang, Miska M. Hannuksela, Viktor Varsa, Ari Hourunranta, Moncef Gabbouj
ICIP (2)1
2000 Block truncation coding with adaptive decimation and interpolation
Ye-Kui Wang, Guofang Tu 0001
VCIP1