VLDB 2026 Research / reviewers in the wild / expert
Gary J. Sullivan
dblp:77/3125
· DBLP profile ↗
48ranked-venue papers
14as first author
9since 2021 · last 2026
0000-0002-2350-9235ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 42 · 12 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Standardizing Generative Face Video Compression Using Supplemental Enhancement InformationabstractThis paper proposes a Generative Face Video Compression (GFVC) approach using Supplemental Enhancement Information (SEI), where a series of compact spatial and temporal representations of a face video signal (e.g., 2D/3D key-points, facial semantics and compact features) can be coded using SEI messages and inserted into the coded video bitstream. At the time of writing, the proposed GFVC approach using SEI messages has been included into a draft amendment of the Versatile Supplemental Enhancement Information (VSEI) standard by the Joint Video Experts Team (JVET) of ISO/IEC JTC 1/SC 29 and ITU-T SG21, which will be standardized as a new version of ITU-T H.274$|$ISO/IEC 23002-7. To the best of the authors' knowledge, the JVET work on the proposed SEI-based GFVC approach is the first standardization activity for generative video compression. The proposed SEI approach has not only advanced the reconstruction quality of early-day Model-Based Coding (MBC) via the state-of-the-art generative technique, but also established a new SEI definition for future GFVC applications and deployment. Experimental results illustrate that the proposed SEI-based GFVC approach can achieve remarkable rate-distortion performance compared with the latest Versatile Video Coding (VVC) standard, whilst also potentially enabling a wide variety of functionalities including user-specified animation/filtering and metaverse-related applications. Yan Ye 0003, Jie Chen 0006, Ru-Ling Liao, Shanzhi Yin, Shiqi Wang 0001, Kaifa Yang, Yue Li 0015, Yiling Xu, Ye-Kui Wang, Shiv Gehlot, Guan-Ming Su, Peng Yin 0002, Sean McCarthy, Gary J. Sullivan |
IEEE Trans. Multim. | 15 |
| 2025 | A Generative Face Video Coding Framework with Disentangled and Consistent BackgroundabstractExisting Generative face video coding (GFVC) frameworks enable the ultra-low bandwidth video communication through transmission of compact facial representations. However, non-localization of derived facial representations leads to foreground and background blending (entanglement) in the decoded sequences, resulting in geometry distortions. In this work, we propose a GFVC framework that removes this blending, and suppresses the induced artifacts. To achieve the disentanglement, the proposed framework 1) separately transmits foreground and background of the first frame (base pictures), and 2) performs fusion at the decoder end with background base picture. Further, the proposed methodology supports chroma keying for decoder simplification and background customization. The proposed approach is generic and builds upon existing GFVC frameworks to generate stable and consistent video sequences. Compared to VVC, proposed algorithm reduces the average bit rate by 52.13% and offers a 1.24% average improvement in DISTS over existing GFVC methods at QP 22. Additionally, subjective evaluations reveal a 88.89% preference for the proposed approach, in contrast to 11.11% for current best-performing method. Shiv Gehlot, Guan-Ming Su, Peng Yin 0002, Sean McCarthy, Gary J. Sullivan |
ICIP | 5 |
| 2024 | The Multiplane Image Information SEI Message and its Use for Distribution of Volumetric Video with Conventional CodecsabstractThis paper describes the background, design and application of a new SEI message – the Multiplane Image Information SEI message, which has recently been adopted into the Technology under Consideration (TuC) document of the JVET committee for potential inclusion in the VSEI standard (ITU-T H.274 and ISO/IEC 23002-7). The paper also provides preliminary compression experiment results and analysis on the implications of the coding efficiency and functionality of the different packing options supported in the SEI message. Taoran Lu, Peng Yin 0002, Guan-Ming Su, Dae Yeol Lee, Tsung-Wei Huang, Sejin Oh, Sean McCarthy, Walt Husak, Gary J. Sullivan |
DCC | 9 |
| 2024 | The source picture timing SEI message in the VSEI standardabstractThe source picture timing information (SPTI) supplemental enhancement information (SEI) message indicates to a receiver the actual temporal distance between source pictures in a manner independent from the output timing of corresponding decoded pictures. The SPTI SEI message can also be used to indicate that a coded picture corresponds to a synthetically generated picture rather than an original source picture. The SPTI SEI message has been agreed to be included in the next version of the versatile supplemental enhancement information (VSEI) standard for use with various video coding standards including VVC, HEVC, and AVC. In this paper, the history, syntax, semantics, and several use cases of the SPTI SEI message are described. Example use cases include high-speed imaging, variable-rate source pictures, event-driven image capture, time-lapse video, time-reverse video, frame rate conversion, motion analysis, and machine analysis for scientific, industrial, and forensic applications. Sean McCarthy, Gary J. Sullivan, Peng Yin 0002 |
VCIP | 2 |
| 2021 | Developments in International Video Coding Standardization After AVC, With an Overview of Versatile Video Coding (VVC)abstractIn the last 17 years, since the finalization of the first version of the now-dominant H.264/Moving Picture Experts Group-4 (MPEG-4) Advanced Video Coding (AVC) standard in 2003, two major new generations of video coding standards have been developed. These include the standards known as High Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC). HEVC was finalized in 2013, repeating the ten-year cycle time set by its predecessor and providing about 50% bit-rate reduction over AVC. The cycle was shortened by three years for the VVC project, which was finalized in July 2020, yet again achieving about a 50% bit-rate reduction over its predecessor (HEVC). This article summarizes these developments in video coding standardization after AVC. It especially focuses on providing an overview of the first version of VVC, including comparisons against HEVC. Besides further advances in hybrid video compression, as in previous development cycles, the broad versatility of the application domain that is highlighted in the title of VVC is explained. Included in VVC is the support for a wide range of applications beyond the typical standard- and high-definition camera-captured content codings, including features to support computer-generated/screen content, high dynamic range content, multilayer and multiview coding, and support for immersive media such as 360° video. Benjamin Bross, Jianle Chen, Jens-Rainer Ohm, Gary J. Sullivan, Ye-Kui Wang |
Proc. IEEE | 4 |
| 2021 | Guest Editorial Introduction to the Special Section on the VVC StandardabstractIn this Special Section of the IEEE Transactions on Circuits and Systems for Video Technology, it is our honor to introduce the Versatile Video Coding (VVC) standard, the latest of the historic partnership collaborations between the International Telecommunication Union Telecommunication Standardization Sector (ITU-T), the International Organization for Standardization (ISO), and the International Electrotechnical Commission (IEC) in the field of video coding standardization. Jill M. Boyce, Jianle Chen, Shan Liu 0001, Jens-Rainer Ohm, Gary J. Sullivan, Thomas Wiegand 0001, Yan Ye 0003, Wenwu Zhu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Overview of the Versatile Video Coding (VVC) Standard and its ApplicationsabstractVersatile Video Coding (VVC) was finalized in July 2020 as the most recent international video coding standard. It was developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO/IEC Moving Picture Experts Group (MPEG) to serve an ever-growing need for improved video compression as well as to support a wider variety of today’s media content and emerging applications. This paper provides an overview of the novel technical features for new applications and the core compression technologies for achieving significant bit rate reductions in the neighborhood of 50% over its predecessor for equal video quality, the High Efficiency Video Coding (HEVC) standard, and 75% over the currently most-used format, the Advanced Video Coding (AVC) standard. It is explained how these new features in VVC provide greater versatility for applications. Highlighted applications include video with resolutions beyond standard- and high-definition, video with high dynamic range and wide color gamut, adaptive streaming with resolution changes, computer-generated and screen-captured video, ultralow-delay streaming, 360° immersive video, and multilayer coding e.g., for scalability. Furthermore, early implementations are presented to show that the new VVC standard is implementable and ready for real-world deployment. Benjamin Bross, Ye-Kui Wang, Yan Ye 0003, Shan Liu 0001, Jianle Chen, Gary J. Sullivan, Jens-Rainer Ohm |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | Overview of the Screen Content Support in VVC: Applications, Coding Tools, and PerformanceabstractIn an increasingly connected world, consumer video experiences have diversified away from traditional broadcast video into new applications with increased use of non-camera-captured content such as computer screen desktop recordings or animations created by computer rendering, collectively referred to as screen content. There has also been increased use of graphics and character content that is rendered and mixed or overlaid together with camera-generated content. The emerging Versatile Video Coding (VVC) standard, in its first version, addresses this market change by the specification of low-level coding tools suitable for screen content. This is in contrast to its predecessor, the High Efficiency Video Coding (HEVC) standard, where highly efficient screen content support is only available in extension profiles of its version 4. This paper describes the screen content support and the five main low-level screen content coding tools in VVC: transform skip residual coding (TSRC), block-based differential pulse-code modulation (BDPCM), intra block copy (IBC), adaptive color transform (ACT), and the palette mode. The specification of these coding tools in the first version of VVC enables the VVC reference software implementation (VTM) to achieve average bit-rate savings of about 41% to 61% relative to the HEVC test model (HM) reference software implementation using the Main 10 profile for 4:2:0 screen content test sequences. Compared to the HM using the Screen-Extended Main 10 profile and the same 4:2:0 test sequences, the VTM provides about 19% to 25% bit-rate savings. The same comparison with 4:4:4 test sequences revealed bit-rate savings of about 13% to 27% for$Y'C_{B}C_{R}$and of about 6% to 14% for$R'G'B'$screen content. Relative to the HM without the HEVC version 4 screen content coding extensions, the bit-rate savings for 4:4:4 test sequences are about 33% to 64% for$Y'C_{B}C_{R}$and 43% to 66% for$R'G'B'$screen content. Tung Nguyen 0001, Xiaozhong Xu, Félix Henry, Ru-Ling Liao, Mohammed Golam Sarwer, Marta Karczewicz, Yung Hsuan Chao, Jizheng Xu, Shan Liu 0001, Detlev Marpe, Gary J. Sullivan |
IEEE Trans. Circuits Syst. Video Technol. | 11 |
| 2021 | The High-Level Syntax of the Versatile Video Coding (VVC) StandardabstractVersatile Video Coding (VVC), a.k.a. ITU-T H.266 | ISO/IEC 23090-3, is the new generation video coding standard that has just been finalized by the Joint Video Experts Team (JVET) of ITU-T VCEG and ISO/IEC MPEG at its$19^{\mathrm {th}}$meeting ending on July 1, 2020. This paper gives an overview of the VVC high-level syntax (HLS), which forms its system and transport interface. Comparisons to the HLS designs in High Efficiency Video Coding (HEVC) and Advanced Video Coding (AVC), the previous major video coding standards, are included. When discussing new HLS features introduced into VVC or differences relative to HEVC and AVC, the reasoning behind the design differences and the benefits they bring are described. The HLS of VVC enables newer and more versatile use cases such as video region extraction, composition and merging of content from multiple coded video bitstreams, and viewport-adaptive 360° immersive media. Ye-Kui Wang, Robert Skupin, Miska M. Hannuksela, Sachin Deshpande, Hendry, Virginie Drugeon, Rickard Sjöberg, Byeongdoo Choi, Vadim Seregin, Yago Sánchez de la Fuente, Jill M. Boyce, Wade Wan, Gary J. Sullivan |
IEEE Trans. Circuits Syst. Video Technol. | 13 |
| 2020 | Versatile Video Coding (VVC) ArrivesabstractSeven years after the development of the first version of the High Efficiency Video Coding (HEVC) standard, the major international organizations in the world of video coding have completed the next major generation, called Versatile Video Coding (VVC). The VVC standard, formally designated as ITU-T H.266 and ISO/IEC 23090-3, promises a major improvement in video compression relative to its predecessors. It can offer roughly double the coding efficiency - i.e., it can be used to encode video content to the same level of visual quality while using about 50% fewer bits than HEVC and thus using about 75% fewer bits than H.264/AVC, today's most widely used format. Thus it can ease the burden on worldwide networks, where video now comprises about 80% of all internet traffic. Moreover, VVC has enhanced features in its syntax for supporting an unprecedented breadth of applications, giving meaning to the word "versatility" used in its title. Completed in July 2020, VVC has begun to emerge in practical implementations and is undergoing testing to characterize its subjective performance. Gary J. Sullivan |
VCIP | 1 |
| 2020 | Guest Editorial Introduction to the Special Section on the Joint Call for Proposals on Video Compression With Capability Beyond HEVCabstractStandardization for digital video compression has shown significant evolution over the last three decades. Starting in 1988 with ITU-T H.261 as the first such standard that was practical for consumer use, ISO/IEC MPEG-1 and H.262/MPEG-2 video (the latter jointly standardized by ITU-T and ISO/IEC) were developed very soon thereafter, creating the first wave of broad usage of digital technology in consumer video, such as broadcast and disc player applications. Later, the H.264/MPEG-4 Advanced Video Coding (AVC) standard was again developed jointly by ITU-T and ISO/IEC experts, with its High Profile becoming dominant from 2004 in HD broadcast and storage, as well as network-based streaming services and private capture of video. With ever-increasing demands for higher quality and the advent of flat-panel displays, the H.265/MPEG-H High Efficiency Video Coding (HEVC) standard became the next generation of video compression standard; with its first version defined in 2013, HEVC has been especially instrumental for the recent deployment of Ultra High Definition (UHD, a.k.a. 4K video). As time has moved forward, video content has continued to become an increasing presence in our lives, with an ever-growing diversification of usage models and continuing demands for higher quality. For example, flat panels evolved towards support of high dynamic range (HDR) video with a wider color gamut, and new modalities for consuming video have appeared, such as head-mounted displays (HMDs). Jill M. Boyce, Jianle Chen, Jens-Rainer Ohm, Gary J. Sullivan, Thomas Wiegand 0001, Yan Ye 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | General Video Coding Technology in Responses to the Joint Call for Proposals on Video Compression With Capability Beyond HEVCabstractAfter the development of the High-Efficiency Video Coding Standard (HEVC), ITU-T VCEG and ISO/IEC MPEG formed the Joint Video Exploration Team (JVET), which started exploring video coding technology with higher coding efficiency, including development of a Joint Exploration Model (JEM) algorithm and a corresponding software implementation. The technology explored in the last version of the JEM further increases the compression capabilities of the hybrid video coding approach by adding new tools, reaching up to 30% bit rate reduction compared to HEVC based on the Bjøntegaard delta bit rate (BD-rate) metric, and further improvement beyond that in terms of subjective visual quality. This provided enough evidence to issue a joint Call for Proposals (CfP) for a new standardization activity now known as Versatile Video Coding (VVC). All technology proposed in the responses to the CfP was based on the classic block-based hybrid video coding design, extending it by new elements of partitioning, intra- and inter-picture prediction, prediction signal filtering, transforms, quantization/scaling, entropy coding, and in-loop filtering. This article provides an overview of technology that was proposed in the responses to the CfP, with a focus on techniques that were not already explored in the JEM context. Benjamin Bross, Kenneth Andersson, Max Bläser, Virginie Drugeon, Seung-Hwan Kim 0001, Jani Lainema, Shan Liu 0001, Jens-Rainer Ohm, Gary J. Sullivan, Ruoyang Yu |
IEEE Trans. Circuits Syst. Video Technol. | 10 |
| 2020 | The Joint Exploration Model (JEM) for Video Compression With Capability Beyond HEVCabstractThis paper provides an overview of the coding algorithms of the Joint Exploration Model (JEM) for video compression with capability beyond HEVC, which was developed by the Joint Video Exploration Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO/IEC Moving Picture Experts Group (MPEG). The goal of the JEM development and experimentation was to provide evidence that sufficient coding efficiency improvement over the High Efficiency Video Coding (HEVC) standard can be achieved, which would justify the need for a new video coding standard with a compression capability significantly exceeding that of HEVC. The development of the JEM provided an ability to conduct studies toward that goal in a verifiable and collaborative manner and led to the launching of the project to develop the new Versatile Video Coding (VVC) standard. Objective metric gains exceeding 30% were measured for most of the tested high-resolution video content that represents current demanding new applications, and subjective testing using human observers showed even more benefit. Jianle Chen, Marta Karczewicz, Yu-Wen Huang, Kiho Choi, Jens-Rainer Ohm, Gary J. Sullivan |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2016 | Introduction to the Special Issue on HEVC Extensions and Efficient HEVC ImplementationsabstractHigh Efficiency Video Coding (HEVC) is the most recent standard in the series of major video coding standards jointly produced by the ITU-T Video Coding Experts Group (VCEG) and the ISO/IEC Moving Picture Experts Group (MPEG). HEVC was first approved in 2013 in the ITU-T as Recommendation H.265 and in ISO/IEC as International Standard 23008-2, and it offers an unprecedented degree of compression capability for a very wide variety of applications. In the three years since its initial completion, it has been extended in several important ways to further broaden its scope. This special issue on HEVC features two sections: 1) HEVC extensions and 2) efficient HEVC implementations. Jens-Rainer Ohm, Gary J. Sullivan, Vivienne Sze, Thomas Wiegand 0001, Madhukar Budagavi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Video Quality Evaluation Methodology and Verification Testing of HEVC Compression PerformanceabstractThe High Efficiency Video Coding (HEVC) standard (ITU-T H.265 and ISO/IEC 23008-2) has been developed with the main goal of providing significantly improved video compression compared with its predecessors. In order to evaluate this goal, verification tests were conducted by the Joint Collaborative Team on Video Coding of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29. This paper presents the subjective and objective results of a verification test in which the performance of the new standard is compared with its highly successful predecessor, the Advanced Video Coding (AVC) video compression standard (ITU-T H.264 and ISO/IEC 14496-10). The test used video sequences with resolutions ranging from 480p up to ultra-high definition, encoded at various quality levels using the HEVC Main profile and the AVC High profile. In order to provide a clear evaluation, this paper also discusses various aspects for the analysis of the test results. The tests showed that bit rate savings of 59% on average can be achieved by HEVC for the same perceived video quality, which is higher than a bit rate saving of 44% demonstrated with the PSNR objective quality metric. However, it has been shown that the bit rates required to achieve good quality of compressed content, as well as the bit rate savings relative to AVC, are highly dependent on the characteristics of the tested content. Thiow Keng Tan, Rajitha Weerakkody, Marta Mrak, Naeem Ramzan, Vittorio Baroncini, Jens-Rainer Ohm, Gary J. Sullivan |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2014 | Subband Decomposition for High-Resolution Color in HEVC and AVC 4: 2: 0 Video Coding SystemsabstractSummary form only given. High-resolution color video (in 4:4:4 format) is becoming increasingly important for screen content and graphics. We have developed a frame packing arrangement (FPA) scheme that enables transmission of 4:4:4 luma-chroma (YUV or YCoCg) content through conventional subsampled chroma (4:2:0) codecs as an approach for low-cost support of 4:4:4 content. We focus on the "band separation" method initially discussed in [1] and evaluate new scenarios with the 4:4:4 and 4:2:0 versions at similar compression levels. We use the HEVC HM12.0 reference software for evaluating the FPA scheme. We vary the luma QP for the main 4:2:0 representation from 10 to 34, keeping the QP for the main frame's chroma and auxiliary frame's luma and chroma (containing three-fourths of the 4:4:4 chroma data) at the same value as the luma QP. We then compare the final 4:4:4 chroma rate-distortion performance of band separation, direct frame packing, and native HM 4:4:4 encoding. We also show preliminary results for lifting-based band separation (with clipping), which eliminates the inherent rounding errors of other band separation filters and provides better rate-distortion performance. Fig. 1 provides an example for a typical case of the new results showing that both the "band separation" and "lifting-based band separation" approaches have much better performance than previously reported, and are significantly better than the earlier-developed "direct" frame packing approach. Although the rate-distortion performance is inferior to that of native HM 4:4:4 coding, the FPA scheme enables the use of existing 4:2:0 decoding hardware, thus allowing for considerable battery power savings in mobile devices. A more complete discussion and evaluation is available in [2]. Srinath Reddy, Sandeep Kanumuri, Shyam Sadhwani, Gary J. Sullivan, Henrique S. Malvar |
DCC | 5 |
| 2013 | Tunneling High-Resolution Color Content through 4: 2: 0 HEVC and AVC Video Coding SystemsabstractWe present a method to convey high-resolution color (4:4:4) video content through a video coding system designed for chroma-sub sampled (4:2:0) operation. The method operates by packing the samples of a 4:4:4 frame into two frames that are then encoded as if they were ordinary 4:2:0 content. After being received and decoded, the packing process is reversed to recover a 4:4:4 video frame. As 4:2:0 is the most widely supported digital color format, the described scheme provides an effective way of transporting 4:4:4 content through existing mass-market encoders and decoders, for applications such as coding of screen content. The described packing arrangement is designed such that the spatial correspondence and motion vector displacement relationships between the nominally-luma and nominally-chroma components are preserved. The use of this scheme can be indicated by a metadata tag such as the frame packing arrangement supplemental enhancement information (SEI) message defined in the HEVC and AVC (Rec. ITU-T H.264 | ISO/IEC 14496-10) video coding standards. In this context the scheme would operate in a similar manner as is commonly used for packing the two views of stereoscopic 3D video for compatible encoding. The technique can also be extended to transport 4:2:2 video through 4:2:0 systems or 4:4:4 video through 4:2:2 systems. Sandeep Kanumuri, Shyam Sadhwani, Gary J. Sullivan, Henrique S. Malvar |
DCC | 5 |
| 2013 | Complexity Reduction and Performance Improvement for Geometry Partitioning in Video CodingabstractGeometry partitioning for video coding involves establishing a partition line boundary within each block-shaped region and applying motion-compensated prediction to the two sub-regions created by the partition line. This paper presents techniques for enhancing the effectiveness and reducing the complexity of geometry partitioning schemes. A texture-difference-based approach is described to simplify the process of selecting the partition lines. Applying this approach together with a described skipping strategy for blocks with uniform texture can achieve a 94% reduction of encoding time while retaining a similar rate-distortion (R-D) performance to the full-search partitioning approach, when implemented for wedge-based geometry partitioning (WGP) in the context of H.264/MPEG-4 AVC JM 16.2. A bit rate improvement of approximately 6% is shown relative to not using geometry partitioning. For further R-D improvement, we describe a background-compensated prediction scheme to reduce the number of overhead bits used for motion vectors. Additionally, for systems in which high-quality depth maps are available, we incorporate depth map usage into the described approaches to generate a more accurate partitioning. Using these approaches with object-boundary-based geometry partitioning can achieve about 9% bit rate savings relative to using WGP, while keeping a similar computational complexity to the described complexity-reduced WGP. Qifei Wang, Xiangyang Ji, Ming-Ting Sun, Gary J. Sullivan, Qionghai Dai |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2012 | Progressive-to-Lossless Compression of Color-Filter-Array Images Using Macropixel Spectral-Spatial TransformationabstractWe present a low-complexity integer-reversible spectral-spatial transform that allows for efficient loss less and lossy compression of color-filter-array images (also referred to as camera-raw images). The main advantage of this new transform is that it maps the pixel array values into a format that can be directly compressed in a loss less, lossy, or progressive-to-loss less manner by an existing typical image coder such as JPEG 2000 or JPEG XR. Thus, no special codec design is needed for compressing the camera-raw data. Another advantage is that the new transform allows for mild compression of camera-raw data in a near-loss less format, allowing for very high quality offline post-processing, but with camera-raw files that can be half the size of those of existing camera-raw formats. Henrique S. Malvar, Gary J. Sullivan |
DCC | 2 |
| 2012 | Compression performance of high efficiency video coding (HEVC) working draft 4abstractThis paper presents the results of compression comparison tests between the current state of the emerging High Efficiency Video Coding (HEVC) draft standard and the current dominant standard H.264/MPEG-4 AVC (High Profile) as an anchor reference. The conditions used for the comparison tests were designed to reflect relevant application scenarios and to enable a fair comparison to the maximum extent feasible, i.e. using comparable quantization settings, reference frame buffering, etc. The testing was generally configured in favour of using a relatively strong H.264/MPEG-4 AVC anchor reference. Several of the encoder optimizations currently found in the HEVC software are tested and shown to be helpful to improve the H.264/MPEG-4 AVC anchor performance. When compared to the improved anchor encoder configurations, the HEVC draft design currently provides a bit rate savings for equal PSNR of about 39% for random access applications, 44% for low-delay use, and 25% for all-intra use. Bin Li 0012, Gary J. Sullivan, Jizheng Xu |
ISCAS | 2 |
| 2012 | Complexity-reduced geometry partition search and high efficiency prediction for video codingabstractTo reduce the complexity of searching for wedge-based geometry partitions in video coding, we propose a texture-difference based partition line selection approach with a skipping strategy. Applying this approach can reduce the encoding time by 90% while retaining the similar rate-distortion performance to that of exhaustive searching. We also propose a background-compensated prediction scheme to improve the rate-distortion performance of the geometry partitioning prediction by reducing its motion vector overhead. Incorporating our proposed approach into object-boundary-based geometry partitioning can achieve about 10% bit-rate savings relative to the full-search approach while keeping the complexity at about the same level as our proposed complexity-reduced wedge-based geometry partitioning. Qifei Wang, Ming-Ting Sun, Gary J. Sullivan |
ISCAS | 3 |
| 2012 | Comparison of the Coding Efficiency of Video Coding Standards - Including High Efficiency Video Coding (HEVC)abstractThe compression capability of several generations of video coding standards is compared by means of peak signal-to-noise ratio (PSNR) and subjective testing results. A unified approach is applied to the analysis of designs, including H.262/MPEG-2 Video, H.263, MPEG-4 Visual, H.264/MPEG-4 Advanced Video Coding (AVC), and High Efficiency Video Coding (HEVC). The results of subjective tests for WVGA and HD sequences indicate that HEVC encoders can achieve equivalent subjective reproduction quality as encoders that conform to H.264/MPEG-4 AVC when using approximately 50% less bit rate on average. The HEVC design is shown to be especially effective for low bit rates, high-resolution video content, and low-delay communication applications. The measured subjective improvement somewhat exceeds the improvement measured by the PSNR metric. Jens-Rainer Ohm, Gary J. Sullivan, Heiko Schwarz, Thiow Keng Tan, Thomas Wiegand 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2012 | Overview of the High Efficiency Video Coding (HEVC) StandardabstractHigh Efficiency Video Coding (HEVC) is currently being prepared as the newest video coding standard of the ITU-T Video Coding Experts Group and the ISO/IEC Moving Picture Experts Group. The main goal of the HEVC standardization effort is to enable significantly improved compression performance relative to existing standards-in the range of 50% bit-rate reduction for equal perceptual video quality. This paper provides an overview of the technical features and characteristics of the HEVC standard. Gary J. Sullivan, Jens-Rainer Ohm, Woojin Han 0001, Thomas Wiegand 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2011 | Reduced-complexity search for video coding geometry partitions using texture and depth dataabstractIn this paper, a texture-space geometry partitioning approach is proposed to reduce the computational complexity of searching for geometry partitions for video coding. Additionally, for systems that capture both video and depth data, an enhanced geometry partitioning approach using both texture and depth information is proposed to further improve the partitioning accuracy and reduce the search complexity. Compared with a full-search approach, the proposed geometry partition search approaches achieve about 94% reduction of the encoding time while retaining similar rate-distortion performance. Qifei Wang, Gary J. Sullivan, Ming-Ting Sun |
VCIP | 3 |
| 2011 | Overview of the Stereo and Multiview Video Coding Extensions of the H.264/MPEG-4 AVC StandardabstractSignificant improvements in video compression capability have been demonstrated with the introduction of the H.264/MPEG-4 advanced video coding (AVC) standard. Since developing this standard, the Joint Video Team of the ITU-T Video Coding Experts Group (VCEG) and the ISO/IEC Moving Picture Experts Group (MPEG) has also standardized an extension of that technology that is referred to as multiview video coding (MVC). MVC provides a compact representation for multiple views of a video scene, such as multiple synchronized video cameras. Stereo-paired video for 3-D viewing is an important special case of MVC. The standard enables inter-view prediction to improve compression capability, as well as supporting ordinary temporal and spatial prediction. It also supports backward compatibility with existing legacy systems by structuring the MVC bitstream to include a compatible “base view.” Each other view is encoded at the same picture resolution as the base view. In recognition of its high-quality encoding capability and support for backward compatibility, the stereo high profile of the MVC extension was selected by the Blu-Ray Disc Association as the coding format for 3-D video with high-definition resolution. This paper provides an overview of the algorithmic design used for extending H.264/MPEG-4 AVC towards MVC. The basic approach of MVC for enabling inter-view prediction and view scalability in the context of H.264/MPEG-4 AVC is reviewed. Related supplemental enhancement information (SEI) metadata is also described. Various “frame compatible” approaches for support of stereo-view video as an alternative to MVC are also discussed. A summary of the coding performance achieved by MVC for both stereo- and multiview video is also provided. Future directions and challenges related to 3-D video are also briefly discussed. Anthony Vetro, Thomas Wiegand 0001, Gary J. Sullivan |
Proc. IEEE | 3 |
| 2010 | Special Section on the Joint Call for Proposals on High Efficiency Video Coding (HEVC) StandardizationabstractThe five papers in this special section were among those submitted in response to the joint call for proposals on high efficiency video coding (HEVC) standardization. Although at this point of development it is still unclear which specific elements the final HEVC standard will contain, the selection of the papers was made such that together they would cover most of the promising tools and technologies that seem likely to be included in the standard. Thomas Wiegand 0001, Jens-Rainer Ohm, Gary J. Sullivan, Woojin Han 0001, Rajan L. Joshi, Thiow Keng Tan, Kemal Ugur |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2007 | New Standardized Extensions of MPEG4-AVC/H.264 for Professional-Quality Video ApplicationsabstractTo support high quality video applications, the Joint Video Team (JVT) has recently added five new profiles, two new supplemental enhancement information (SEI) messages, and two new extended gamut color space indicators to the MPEG4-AVC/H.264 video coding standard. The new profiles include substantial feature enhancements for high-quality video applications, including improved-efficiency 4:4:4 video format coding, improved-efficiency lossless macroblock coding, coding 4:4:4 video pictures using three separately-coded color planes, and support of bit depths up to 14 bits per sample. The new features were developed to support a wide range of applications where high quality video compression is demanded, including professional and semi-professional scenarios in particular. They also anticipate the introduction of higher fidelity displays. In this paper, the new extensions are presented along with quantitaive estimates of the benefits of the new features and a discussion of the target application environments. Gary J. Sullivan, Haoping Yu, Shun-ichi Sekiguchi, Huifang Sun, Thomas Wedi, Steffen Wittmann, Yung Lyul Lee, C. Andrew Segall, Teruhiko Suzuki |
ICIP (1) | 1 |
| 2007 | Spatial Scalability Within the H.264/AVC Scalable Video Coding ExtensionabstractA scalable extension to the H.264/AVC video coding standard has been developed within the joint video team (JVT), a joint organization of the ITU-T video coding group (VCEG) and the ISO/IEC moving picture experts group (MPEG). The extension allows multiple resolutions of an image sequence to be contained in a single bit stream. In this paper, we introduce the spatially scalable extension within the resulting scalable video coding standard. The high-level design is described and individual coding tools are explained. Additionally, encoder issues are identified. Finally, the performance of the design is reported. C. Andrew Segall, Gary J. Sullivan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Introduction to the Special Issue on Scalable Video Coding-Standardization and BeyondabstractThe thirteen papers in this special issue are devoted to the standardization and development of scalable video coding techniques and applications. Thomas Wiegand 0001, Gary J. Sullivan, Jens-Rainer Ohm, Ajay Luthra |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | Improved lossless intra coding for H.264/MPEG-4 AVCabstractA new lossless intra coding method based on sample-by-sample differential pulse code modulation (DPCM) is presented as an enhancement of the H.264/MPEG-4 AVC standard. The H.264/AVC design includes a multidirectional spatial prediction method to reduce spatial redundancy by using neighboring samples as a prediction for the samples in a block of data to be encoded. In the new lossless intra coding method, the spatial prediction is performed based on samplewise DPCM instead of in the block-based manner used in the current H.264/AVC standard, while the block structure is retained for the residual difference entropy coding process. We show that the new method, based on samplewise DPCM, does not have a major complexity penalty, despite its apparent pipeline dependencies. Experiments show that the new lossless intra coding method reduces the bit rate by approximately 12% in comparison with the lossless intra coding method previously included in the H.264/AVC standard. As a result, the new method is currently being adopted into the H.264/AVC standard in a new enhancement project. Yung Lyul Lee, Ki-Hun Han, Gary J. Sullivan |
IEEE Trans. Image Process. | 3 |
| 2005 | Video Compression - From Concepts to the H.264/AVC StandardabstractOver the last one and a half decades, digital video compression technologies have become an integral part of the way we create, communicate, and consume visual information. In this paper, techniques for video compression are reviewed, starting from basic concepts. The rate-distortion performance of modern video compression schemes is the result of an interaction between motion representation techniques, intra-picture prediction techniques, waveform coding of differences, and waveform coding of various refreshed regions. The paper starts with an explanation of the basic concepts of video codec design and then explains how these various features have been integrated into international standards, up to and including the most recent such standard, known as H.264/AVC. Gary J. Sullivan, Thomas Wiegand 0001 |
Proc. IEEE | 1 |
| 2004 | On embedded scalar quantizationabstractThe paper studies the rate-distortion performance of symmetric scalar quantizers having a large (effectively infinite) number of steps and using the same step size for all steps except the one containing the zero input value. Quantizers of this form have been shown to have good performance for a variety of sources, and are precisely optimal for the Laplacian source. The performance is investigated particularly for embedded quantization, in which the representation of a source quantity is refined successively by forming finer quantizers from further segmentation of the steps of coarser quantizer constructions. Although the use of a double-wide dead-zone has dominated prior embedded quantization practice, it is shown that any rational number can be maintained as a stable dead-zone ratio. Two forms are investigated in more depth - quantizers with dead-zone ratios of 1 and 2 - and a ratio of 1 is shown often to provide a significant performance advantage (up to 1 dB). Performance is explored primarily in the context of the generalized Gaussian pdf using the squared-error distortion measure, but should also apply in other contexts. Gary J. Sullivan |
ICASSP (4) | 1 |
| 2003 | Introduction to the special issue on the H.264/AVC video coding standardabstractS.557-559 Gary J. Sullivan, Ajay Luthra, Thomas Wiegand 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2003 | Overview of the H.264/AVC video coding standardabstractH.264/AVC is newest video coding standard of the ITU-T Video Coding Experts Group and the ISO/IEC Moving Picture Experts Group. The main goals of the H.264/AVC standardization effort have been enhanced compression performance and provision of a "network-friendly" video representation addressing "conversational" (video telephony) and "nonconversational" (storage, broadcast, or streaming) applications. H.264/AVC has achieved a significant improvement in rate-distortion efficiency relative to existing standards. This article provides an overview of the technical features of H.264/AVC, describes profiles and applications for the standard, and outlines the history of the standardization process. Thomas Wiegand 0001, Gary J. Sullivan, Gisle Bjøntegaard, Ajay Luthra |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2003 | Rate-constrained coder control and comparison of video coding standardsabstractA unified approach to the coder control of video coding standards such as MPEG-2, H.263, MPEG-4, and the draft video coding standard H.264/AVC (advanced video coding) is presented. The performance of the various standards is compared by means of PSNR and subjective testing results. The results indicate that H.264/AVC compliant encoders typically achieve essentially the same reproduction quality as encoders that are compliant with the previous standards while typically requiring 60% or less of the bit rate. Thomas Wiegand 0001, Heiko Schwarz, Anthony Joch, Faouzi Kossentini, Gary J. Sullivan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2002 | Performance comparison of video coding standards using Lagrangian coder controlabstractA unified approach to the coder control of video coding standards such as MPEG-2, H.263, MPEG-4, and the draft video coding standard JVT/H.26L/AVC is presented. Using this unified framework, the performance of the various standards is compared by means of PSNR and subjective testing results. The results indicate that JVT/H.26L/AVC compliant encoding can typically achieve essentially the same objective PSNR reproduction quality as encoders that are compliant with previous standards while requiring as little as 60% or less of the bit rate of the next best standard, particularly for higher-latency applications and particularly for more difficult source material. Subjective testing shows that the bit savings produced by this draft standard are even larger than the PSNR results indicate. Anthony Joch, Faouzi Kossentini, Heiko Schwarz, Thomas Wiegand 0001, Gary J. Sullivan |
ICIP (2) | 5 |
| 1996 | Lapped Orthogonal Vector QuantizationabstractThe blocking artifacts that arise in the use of traditional vector quantization (VQ) schemes can, in general, be virtually eliminated via an efficient lapped VQ strategy. With lapped VQ, blocks are obtained from the source in an overlapped manner, and reconstructed via superposition of overlapped codevectors. The new scheme, which we term lapped orthogonal vector quantization (LOVQ), requires no increase in bit rate and, in contrast to other proposed approaches, no significant increase in computational complexity or memory requirements. Attractively, the use of LOVQ also leads to a modest increase in coding gain over traditional VQ schemes of comparable complexity. Henrique S. Malvar, Gary J. Sullivan, Gregory W. Wornell |
Data Compression Conference | 2 |
| 1996 | Efficient scalar quantization of exponential and Laplacian random variablesabstractThis paper presents solutions to the entropy-constrained scalar quantizer (ECSQ) design problem for two sources commonly encountered in image and speech compression applications: sources having the exponential and Laplacian probability density functions. We use the memoryless property of the exponential distribution to develop a new noniterative algorithm for obtaining the optimal quantizer design. We show how to obtain the optimal ECSQ either with or without an additional constraint on the number of levels in the quantizer. In contrast to prior methods, which require a multidimensional iterative solution of a large number of nonlinear equations, the new method needs only a single sequence of solutions to one-dimensional nonlinear equations (in some Laplacian cases, one additional two-dimensional solution is needed). As a result, the new method is orders of magnitude faster than prior ones. We show that as the constraint on the number of levels in the quantizer is relaxed, the optimal ECSQ becomes a uniform threshold quantizer (UTQ) for exponential, but not for Laplacian sources. We then further examine the performance of the UTQ and optimal ECSQ, and also investigate some interesting alternatives to the UTQ, including a uniform-reconstruction quantizer (URQ) and a constant dead-zone ratio quantizer (CDZRQ). Gary J. Sullivan |
IEEE Trans. Inf. Theory | 1 |
| 1994 | Optimal entropy constrained scalar quantization for exponential and Laplacian random variablesabstractThis paper presents solutions to the entropy-constrained scalar quantizer (ECSQ) design problem for two sources commonly encountered in image and speech compression applications: sources having exponential and Laplacian probability density functions. We obtain the optimal ECSQ either with or without an additional constraint on the number of levels in the quantizer. In contrast to prior methods, which require iterative solution of a large number of nonlinear equations, the new method needs only a single sequence of solutions to one-dimensional nonlinear equations (in some Laplacian cases, one additional two-dimensional solution is needed). As a result, the new method is orders of magnitude faster than prior ones. We also show that as the constraint on the number of levels in the quantizer is relaxed, the optimal ECSQ becomes a uniform threshold quantizer (UTQ) for exponential, but not for Laplacian sources.> Gary J. Sullivan |
ICASSP (5) | 1 |
| 1994 | Methods of Reduced-Complexity Overlapped Block Motion CompensationabstractOverlapped block motion compensation (OBMC) can significantly improve upon the prediction performance of conventional block motion compensation, though at the cost of increased decoder computational complexity. This paper analyzes the complexity costs of OBMC, and offers several modifications to OBMC that significantly reduce its decoder complexity while retaining most of the OBMC performance gain.> Gary J. Sullivan, Michael T. Orchard |
ICIP (2) | 1 |
| 1994 | Overlapped block motion compensation: an estimation-theoretic approachabstractWe present an estimation-theoretic analysis of motion compensation that, when used with fields of block-based motion vectors, leads to the development of overlapped block algorithms with improved compensation accuracy. Overlapped block motion compensation (OBMC) is formulated as a probabilistic linear estimator of pixel intensities given the limited block motion information available to the decoder. Although overlapped techniques have been observed to reduce blocking artifacts in video coding, this analysis establishes for the first time how (and why) OBMC can offer substantial reductions in prediction error as well, even with no change in the encoder's search and no extra side information. Performance can be further enhanced with the use of state variable conditioning in the compensation process. We describe the design of optimized windows for OBMC. We also demonstrate how, with additional encoder complexity, a motion estimation algorithm optimized for OBMC offers further significant gains in compensation accuracy. Overall mean-square prediction improvements in the range of 16 to 40% (0.8 to 2.2 dB) are demonstrated. Michael T. Orchard, Gary J. Sullivan |
IEEE Trans. Image Process. | 2 |
| 1994 | Efficient quadtree coding of images and videoabstractThe quadtree data structure is commonly used in image coding to decompose an image into separate spatial regions to adaptively identify the type of quantizer used in various regions of an image. The authors describe the theory needed to construct quadtree data structures that optimally allocate rate, given a set of quantizers. A Lagrange multiplier method finds these optimal rate allocations with no monotonicity restrictions. They use the theory to derive a new quadtree construction method that uses a stepwise search to find the overall optimal quadtree structure. The search can be driven with either actual measured quantizer performance or ensemble average predicted performance. They apply this theory to the design of a motion compensated interframe video coding system using a quadtree with vector quantization. Gary J. Sullivan, Richard L. Baker |
IEEE Trans. Image Process. | 1 |
| 1993 | Multi-hypothesis motion compensation for low bit-rate video coding
Gary J. Sullivan |
ICASSP (5) | 1 |
| 1992 | Rate-distortion optimization for tree-structured source coding with multi-way node decisionsabstractA new algorithm is proposed for the generation of rate-distortion optimized tree-structures used in source coding. The algorithm solves a more general problem than that assumed for a previous method by P.A. Chou et al. (1989). It arises when the tree structure is to be optimized with a multi-way decision made at each leaf node. The simplifying assumption that the three structure performance functions are monotone with respect to the level of nodes in the tree is also removed. Two applications of the more general problem in image/video compression are quadtree coding with vector quantization and variable block-size motion compensation. This new method is not a pruning algorithm, as the optimal trees may not be nested. The new algorithm finds the global optimal tree structure and node decisions in the unconstrained rate allocation sense.> Gary J. Sullivan, Richard L. Baker |
ICASSP | 1 |
| 1992 | Recursive optimal pruning with applications to tree structured vector quantizersabstractA pruning algorithm of P.A. Chou et al. (1989) for designing optimal tree structures identifies only those codebooks which lie on the convex hull of the original codebook's operational distortion rate function. The authors introduce a modified version of the original algorithm, which identifies a large number of codebooks having minimum average distortion, under the constraint that, in each step, only modes having no descendents are removed from the tree. All codebooks generated by the original algorithm are also generated by this algorithm. The new algorithm generates a much larger number of codebooks in the middle- and low-rate regions. The additional codebooks permit operation near the codebook's operational distortion rate function without time sharing by choosing from the increased number of available bit rates. Despite the statistical mismatch which occurs when coding data outside the training sequence, these pruned codebooks retain their performance advantage over full search vector quantizers (VQs) for a large range of rates. Shei-Zein Kiang, Richard L. Baker, Gary J. Sullivan, Chung-Yen Chiu |
IEEE Trans. Image Process. | 3 |
| 1991 | Recursive optimal pruning of tree-structured vector quantizersabstractThe generalized BFOS (G-BFOS), a sequential pruning algorithm for designing optimal tree structures, was presented by Chou, Lookabaugh, and Gray (see IEEE Trans. Inf. Theory, vol.35, no.2, p.299, 1989), and it was applied to tree-structured vector quantizers (TSVQ). G-BFOS yields VQ codebooks that often outperform conventional generalized Lloyd full search codebooks having the same rate and block size. The authors have developed a modified version of G-BFOS, called the recursive optimal pruning algorithm (ROPA), which recursively searches for the nodes to be pruned next. The sequence of these pruned codebooks includes the original optimal G-BFOS codebooks and many additional ones. The optimality of these codebooks is described, and simulations evaluate their performance.> Shei-Zein Kiang, Gary J. Sullivan, Chung-Yen Chiu, Richard L. Baker |
ICASSP | 2 |
| 1991 | Efficient quadtree coding of images and videoabstractThe authors describe the theory needed to construct quadtree data structures which optimally allocate rate, given a set of quantizers. A Lagrange multiplier method finds these optimal rate allocations with no monotonicity restrictions. The theory is used to derive a new quadtree construction method which uses a stepwise search to find the overall optimal quadtree structure. The search can be driven with either actual measured quantizer performance or ensemble average predicted performance. This theory is then applied to the design of an interframe hybrid video coding system using a quadtree with vector quantization.> Gary J. Sullivan, Richard L. Baker |
ICASSP | 1 |
| 1991 | Motion compensation for video compression using control grid interpolationabstractA new class of motion compensation methods that are based on control grid interpolation (CGI) for use in video compression is described. The predominant motion compensation method, block matching, is shown to be a special case of CGI. A new CGI method is presented that produces a smooth motion vector field, preserving continuity and connectivity in the prediction image. When contrasted with block matching, the new algorithm eliminates blocking artifacts while using the same or less side information for approximately the same mean square error. Search algorithms and coding methods developed previously for block matching can be used with this new technique.> Gary J. Sullivan, Richard L. Baker |
ICASSP | 1 |