C. Andrew Segall

dblp:30/280 · also Andrew Segall · DBLP profile ↗
← Back
48ranked-venue papers
18as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 47 · 18 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2025 Ultra-Low Complexity Neural Networks for Next Generation Video Decoding
abstract
We consider the problem of embedding a neural network directly into a video decoder. This requires a design with complexity suitable for implementation on mobile and power constrained devices. To achieve this goal, we explored Multi-scale CNN (MSCNN) design in [1]. In this paper, we improve the design to support super resolution spatial scale factors SF==(1.5×, 2×, 3×, 4×, 6×) by modifying the polyphase filter (Figure 1a) that generates an upsampled output using g(scale) phases and stride of Sscale. When SF= 1.5 ×, g(scale, Sscale) = (9,2); Otherwise it is (scale2,1). gsG, kK, and sS denote channel group size of G, kernel size of K×K, and stride of S. To reduce per-pixel Multiply-Accumulates (MACs), the 3×1 and 1×3 convolutional layers use Canonical Polyadic (CP) decomposition and reduced channel count. These changes reduce MACs/pixel from 1,924 in [1] to 1,192 to 584. Figure 1b, shows the placement of MSCNN in AVM [2]. We code 4K video, using AOMedia's Adaptive Streaming (AS) test conditions and compare MSCNN versus following resampler combinations: Downsampling - [L5: Lanczos(5), L6: Lanczos(6)]; Upsampling - [L5, L6, BL: Bilinear, BC: Bicubic]. We observe MSCNN provides on average 30.4% rate saving.
Kiran M. Misra, Shashwat Ranjan Chaurasia, C. Andrew Segall, Byeongdoo Choi
DCC3
2025 Efficient Random Access Method Using Seed and Inter-Key Frames for Next Generation Video Codec
abstract
Random access points are a key property of a video coding system. These points indicate where a decoder can start decoding, and they traditionally correspond to frames that are not predicted from previous data. In this paper, we revisit the random access problem in the context of modern video streaming and over-the-top transport systems. We propose that these systems employ an alternative approach that relies on "seed" frames that are periodically provided to the decoder. These frames persist in decoder memory and can be used for prediction of each random access point. Experimental results show the efficacy of the proposed approach. Specifically, we observe a 6.50% reduction in bit-rate when using the AOMedia common test conditions, an 18.38% reduction when emulating live sports events, and a 35.03% improvement for security applications when measured using VMAF.
Byeongdoo Choi, C. Andrew Segall, Kiran M. Misra
ICIP2
2023 Multiscale convolutional neural networks for in-loop video restoration
abstract
Incorporating neural networks into a video codec as an in-loop filter has been shown to provide significant improvements in coding efficiency. Unfortunately, the computational complexity associated with the neural network, specifically the number of multiply-accumulate (MAC) operations, makes these approaches intractable in practice. In this paper, we consider using a multiscale approach to reduce complexity while maintaining coding efficiency. Experimental results demonstrate a 5.4× reduction in MAC operations while achieving an average bit rate savings of 6.4% and 6.3% for all intra and random access coding, respectively, when compared to the evolving AV2 standard. Ablation studies are also provided and show that the approach achieves all but 0.2% of the coding efficiency of full resolution processing.
Kiran M. Misra, C. Andrew Segall, Byeongdoo Choi
DCC2
2023 Reduced Complexity Multiscale CNN for in-Loop Video Restoration
abstract
Convolutional neural networks (CNNs) have shown promising improvements in video coding efficiency when included in traditional block-based codecs as a loop filter. Unfortunately, these coding gains are often accompanied by significant increases in complexity, measured by the number of multiply-accumulate (MAC) operations, that make them intractable in practice. As a result, there is considerable interest in reducing complexity for these CNN-based approaches. In previous work, we have shown that multiscale CNNs provide a path to reduce the associated MAC count. In this paper, we extend our work to consider channel grouping, spatial support limitations and shallower network depths to further reduce the MAC count of these multi-scale architectures. We demonstrate that the method can achieve an average VMAF bitrate reduction of 6.1% and 2.6% for all intra and random-access coding respectively, when compared to the evolving AV2 standard. Complexity is reduced to 1.85k MACs per pixel, which is a 390× reduction over previously published results.
Kiran M. Misra, C. Andrew Segall, Byeongdoo Choi
ICIP2
2022 Video Feature Compression for Machine Tasks
abstract
We consider the problem of transmitting video from a remote device to a cloud-based classification system in a bandwidth limited network. Our focus is on developing an end-to-end system that extracts features from the video data and compresses these features for transmission. In this paper, we consider approaches that operate on each video frame independently as well as exploiting the temporal correlation between frames. In both cases, the transmitted features can be used for object detection and instance segmentation tasks using existing, pre-trained networks. Results show the efficacy of the approach with improvements in coding efficiency ranging from 46.3% to 92.8% when compared to compressing the video data using state-of-the-art video compression standards.
Kiran M. Misra, Tianying Ji, C. Andrew Segall, Frank Bossen
ICME3
2020 High Dynamic Range Video Coding Technology in Responses to the Joint Call for Proposals on Video Compression With Capability Beyond HEVC
abstract
The ITU-T Video Coding Experts Group and the ISO/IEC Moving Picture Experts Group issued a Call for Proposals (CfP) on video compression with capability beyond HEVC in October 2017. The CfP considered three categories of content - Standard Dynamic Range, High Dynamic Range and Wide Colour Gamut (HDR/WCG), and 360° Omni-directional video. As a result of the CfP process, the development of a new video coding standard, named Versatile Video Coding (VVC), was initiated. The goal of this paper is to provide an overview of the CfP responses for the HDR/WCG category. The paper includes a summary of work leading to the development of the CfP, a presentation of the CfP results for the HDR/WCG category, and a description of the specific HDR/WCG technologies submitted to the CfP.
Edouard François, C. Andrew Segall, Alexis M. Tourapis, Peng Yin 0002, Dmytro Rusanovskyy
IEEE Trans. Circuits Syst. Video Technol.2
2020 Tools for Video Coding Beyond HEVC: Flexible Partitioning, Motion Vector Coding, Luma Adaptive Quantization, and Improved Deblocking
abstract
This paper provides a description and analysis of technology included in two contributions to the Call for Proposals for Video Compression with Capability beyond HEVC. This Call for Proposals was issued jointly by the Moving Picture Experts Group (MPEG) and the Video Coding Experts Group (VCEG). The contribution emphasized a flexible and rectangular partitioning structure, which was combined with both new and existing coding tools. New coding tools included methods for improved motion vector coding, quantization signaling, and deblocking; while existing tools were largely methods studied in the Joint Exploration Model software. Results show the efficacy of the approach. Using the sequences and test conditions defined in the Call for Proposals, the described approach provided an average bitrate reduction, relative to an HEVC anchor, of 41.2% and 35.7% for 4K and HD test sequences, respectively. Moreover, the method achieved a compression performance of 34.3% and 32.2% for high dynamic range content using the perceptual quantizer (PQ) and Hybrid-Log Gamma transfer functions, respectively.
Kiran M. Misra, C. Andrew Segall, Frank Bossen
IEEE Trans. Circuits Syst. Video Technol.2
2019 Enhanced Compression beyond HEVC for Next Generation Content
abstract
The Joint Video Experts Team recently evaluated technology in response to a Call for Proposals for Video Compression with Capability beyond HEVC. A number of proposed solutions were evaluated, with a sub-set demonstrating the potential to reduce bit-rates by over 40% compared to HEVC. This paper presents the author's contributions to one of these proposals. The proposal emphasized a flexible, rectangular partitioning structure that was combined with new coding tools, including improved motion vector coding and quantization signaling methods. Results show the efficacy of the approach. Using the evaluation procedure defined in the Call, the described approach provides coding gains relative to an HEVC anchor of 41.2% and 35.7% for 4K-SDR and HD-SDR sequences, respectively, using the random access configuration; 29.0% for HD-SDR sequences using a low delay configuration, and gains of 34.3% and 32.2% for PQ-HDR and HLG-HDR sequences, respectively, using a random access configuration.
Kiran M. Misra, C. Andrew Segall, Weijia Zhu, Byeongdoo Choi, Frank Bossen, Phil Cowan
DCC2
2019 On Cross Component Adaptive Loop Filter for Video Compression
abstract
This paper introduces a video coding tool called the Cross Component Adaptive Loop Filter (CC-ALF). The goal of the tool is to improve chroma fidelity, and this is achieved by exploiting correlations between the luma and chroma channels during the loop filtering process. The tool analyzes the luma channel to determine a chroma residual that is added to a previously reconstructed chroma signal, where the analysis consists of applying a linear filter to the luma information. Adaptation of the filter is achieved by signaling the filter coefficients in the bit-stream and, furthermore, selectively enabling or disabling the process spatially across a picture. Results show the efficacy of the approach, where the performance is evaluated using the Common Test Conditions (CTC) defined by Joint Video Experts Team (JVET). The evaluations show that CCALF provides average chroma Bjøntegaard Delta improvement of 7.6%, 13.8% and 17.8% for all intra, random access, and low delay configurations. Additionally, an approach to shift these chroma gains to luma gains is considered and results in an average luma BD bitrate improvement of 0.9%.
Kiran M. Misra, Frank Bossen, C. Andrew Segall
PCS3
2018 Compound Split Tree for Video Coding
abstract
During the exploration of video coding technology for potential next generation standards, the Joint Video Exploration Team (JVET) has been studying quad-tree plus binary-tree (QTBT) partition structures within its Joint Exploration Model (JEM). This QTBT partition structure provides more flexibility compared with the quad-tree only partition structure in HEVC. Here, we further consider the QTBT structure and extended it to allow quad-tree partitioning to be performed both before and after a binary-tree partition. We refer to this structure as a compound split tree (CST). To show the efficacy of the approach, we implemented the method into JEM7. The method achieved 1.25%, 2.11% and 1.87% BD-bitrate savings for Y, U and V components on average under the random-access configuration, respectively.
Weijia Zhu, C. Andrew Segall
PCS2
2017 Spatially Scalable HEVC for Layered Division Multiplexing in Broadcast
abstract
Recent broadcast standards support Layered Division Multiplexing (LDM) to achieve graceful degradation as signal quality degrades at the receiver. LDM is accomplished by using different constellations within the same Radio Frequency (RF) spectrum. LDM thus enables delivering multiple service tiers in a single broadcast channel. LDM when used in conjunction with scalable source coding codecs such as the Scalable extension of High Efficiency Video Coding (SHVC), further helps improve overall spectrum utilization and efficiency. In this paper we investigate a 2-tier broadcast LDM based service with one service tier aimed at lower video resolution such as 540p, 720p, 1080p for a mobile receiver (smaller/indoor antenna) and the other service tier targeting twice the video resolution of the lower tier, for stationary receivers (larger/outdoor antenna). The primary contribution of this paper is to identify 2-tier transmission configurations of interest to broadcasters and compare the spectrum and bitrate coding efficiency gains of an SHVC-based multi-tier service versus a simulcast (single layer) based multi-tier service for an Advanced Television Systems Committee (ATSC) 3.0 transmission system. Bitrate savings ranging from 38% and 57% is observed for the SHVC based layered system. For large coverage and pedestrian with a receiver test scenarios, channel utilization savings ranging from 23% to 46% is observed. For mobile and tablet in bedroom scenarios a smaller broadcast bandwidth savings ranging from 6% to 9% is observed.
Kiran M. Misra, C. Andrew Segall, Jie Zhao 0007, Seung-Hwan Kim 0001, Joan Llach, Alan Stein, John Stewart, Hendry, Ye-Kui Wang, Yan Ye 0003
DCC2
2016 Quantizer noise equalization between electro-optical transfer functions
abstract
We consider the problem of equalizing the quantization noise between two video encoders operating with video stored using different transfer function. Here, a transfer function is the mapping between linear light and the quantized code words and includes, for example, the well-known “gamma curve” used for most video applications. Our motivation for studying this problem is that the solution enables the direct conversation of a quantization strategy developed for a legacy, standard dynamic range transfer function to be easily mapped to a quantization strategy for a new, high dynamic range transfer function. This is especially timely given the beginning deployment of high dynamic range services throughout the world that use the BT-2020/SMPTE ST 2084 transfer function and not the BT.709/BT.1886 transfer function historically employed for coding high definition digital video. The paper derives a model for the needed relationship. Results show the benefit of the approach.
C. Andrew Segall, Jie Zhao 0007, Seung-Hwan Kim 0001, Kiran M. Misra
PCS1
2016 High quality HDR video compression using HEVC main 10 profile
abstract
This paper describes high-quality compression of high dynamic range (HDR) video using existing tools such as the HEVC Main 10 profile, the SMPTE ST 2084 (PQ) transfer function, and the BT.2020 non-constant luminance Y'CbCr color representation. First, we present novel mathematical bounds that reduce complexity of luminance-preserving subsampling (luma adjustment). A nested look-up table allows for further speedup. Second, an adaptive QP scheme is presented that obtains a better bit allocation balance between dark and bright areas of the picture. Third, a method to control the bit allocation balance between chroma and luma by adjusting the chroma QP offset is presented. The result is a considerable increase in perceptual quality compared to the anchors used in the 2015 MPEG High Dynamic Range/Wide Color Gamut Call for Evidence. All techniques are encoder-side-only, making them compatible with a regular decoder capable of supporting HEVC Main10/PQ/BT.2020, which is already available in some TV sets on the market.
Jacob Ström, Kenneth Andersson, Martin Pettersson, Per Hermansson, Jonatan Samuelsson, C. Andrew Segall, Jie Zhao 0007, Seung-Hwan Kim 0001, Kiran M. Misra, Alexis M. Tourapis, Yeping Su, David Singer
PCS6
2016 High Dynamic Range and Wide Color Gamut Video Coding in HEVC: Status and Potential Future Enhancements
abstract
As the video industry begins deployment of ultrahigh-definition TV in both professional and consumer markets, including support for higher dynamic range and wider color gamut services is considered essential within the industry. Higher dynamic range and wider color gamut offer end users a significantly enhanced viewing experience by supporting intensity ranges and colors unattainable in existing distribution ecosystems. In response to this trend, several standardization organizations have launched efforts to better enable these features in both short term and midterm. In this paper, we provide a survey of these standardization activities, with the specific goal of providing a summary of the underlying technologies. Our emphasis is on both existing and potential extensions to the High Efficiency Video Coding standard.
Edouard François, Chad Fogg, Yuwen He, Xiang Li 0003, Ajay Luthra, C. Andrew Segall
IEEE Trans. Circuits Syst. Video Technol.6
2015 Facial video super resolution using semantic exemplar components
abstract
We present a method for video super resolution using exemplar images of semantic components. In previous work, we proposed a novel super resolution framework based on semantic components and applied it to still images of human faces. In this paper, we extend the approach to video sequences and propose several methods to overcome temporal jitter that results from standard single frame processing. To achieve consistent selection of facial components from a database of exemplars, we introduce a weighted histogram constructed over a temporal window. We then use pixel-based alignment between the exemplar and input image to reduce temporal jitter of the selected component. To further improve temporal stability, we include a temporal constraint into a final optimization stage that blends high resolution exemplar image data into the upscaled input image. We compare our results on face video clips to those of several state-of-the-art super resolution methods, demonstrating the efficacy of the proposed approach.
Anustup Choudhury, Peter van Beek, C. Andrew Segall
ICIP4
2013 Color Gamut Scalable Video Coding
abstract
This paper describes a scalable extension of the High Efficiency Video Coding (HEVC) standard that supports different color gamuts in an enhancement and base layer. Here, the emphasis is on scenarios with BT.2020 color gamut in an enhancement layer and BT.709 color gamut in the base layer. This is motivated by a need to provide content for both high definition and ultra-high definition devices in the near future. The paper describes a method for predicting the enhancement layer samples from a decoded base layer using a series of multiplies and adds to account for both color gamut and bit-depth changes. Results show an improvement in coding efficiency between 65% and 84% for luma (57% and 85% for chroma) compared to simulcast in quality (SNR) scalable coding.
Louis Kerofsky, C. Andrew Segall, Seung-Hwan Kim 0001
DCC2
2013 Kernel smoothing for jagged edge reduction
abstract
In this paper, we consider the problem of removing jaggy artifacts from images. We consider the kernel regression framework and propose a reduced-rank quadratic adaptive method that adapts to the local gradient direction. The proposed technique is effective in shrinking isophote fluctuations, and the result is smooth edges. We observe that it is critical to differentiate jaggy artifacts from texture, junctions and corners, so that meaningful image structure is preserved. Here, we demonstrate that the spectrum of the local covariance matrix of gradients, also known as the structure tensor, is well suited for differentiation of jaggy artifacts from image structure, and we incorporate this into the kernel regression framework. Results show the efficacy of the approach. Namely, that the method is effective in reducing jaggy artifacts without blurring meaningful image structure.
Mohammad Aghagolzadeh, C. Andrew Segall
ICASSP2
2013 Content-adaptive upsampling for scalable video coding
abstract
In the developing scalable extension of the HEVC/H.265 standard, a low-resolution baselayer picture may be used for predicting a higher resolution enhancement layer. This requires an upsampling process to generate the prediction, and this upsampling process is traditionally a linear and time invariant interpolator. In this paper, we consider an upsampling design that is both non-linear and content adaptive. This choice is motivated by the compression noise in the baselayer. We propose a novel approach to include content-aware filtering into the upsampling process. It has low complexity. More importantly, it has low latency and uses the same number of line buffers as a linear interpolator. Results show the efficacy of the method. Specifically, using the test model for the scalable extension of HEVC/H.265, we observe an average bit-rate reduction of 1.1%, when the change in resolution between layers is a factor of two in each dimension.
Jie Zhao 0007, Kiran M. Misra, C. Andrew Segall
PCS3
2012 Image detail enhancement using a dictionary technique
abstract
We present a novel approach to detail enhancement using a dictionary-based technique. For each low-resolution input image patch, we seek a sparse representation from an over-complete dictionary and use that to estimate the high-resolution patch. We modify an existing dictionary-based super-resolution method in several ways to achieve enhancement of fine detail without introduction of new artifacts. These modifications include adaptive enhancement of reconstructed detail patches based on edge analysis to avoid halo artifacts and using an adaptive regularization term to enable noise suppression while enhancing detail. We compare with state-of-the-art methods and show better results in terms of enhancement with suppression of noise.
Anustup Choudhury, Peter van Beek, C. Andrew Segall
ICIP3
2012 On transform dynamic range in high efficiency video coding
abstract
We consider the problem of determining the dynamic range of an inverse transform. An analytic approach is described, and the resulting bound is a function of the bit-depth of the input data and the maximum L1norm of the transformation matrix. Furthermore, the analytic approach is refined to account for nonlinearities within a transform due to intermediate, integer conversions. This refinement considers the maximum error introduced by the nonlinearity on the initial, linear bound. The efficacy of the proposed tool is demonstrated by analyzing the inverse transform in the high efficiency video coding (HEVC) standard developed by the JCT-VC.
Louis Kerofsky, Kiran M. Misra, C. Andrew Segall
ICIP3
2012 Tiles for managing computational complexity of video encoding and decoding
abstract
In this paper, we introduce the concept of tiles. Tiles are incorporated into the current design of the High Efficiency Video Coding (HEVC) standard being developed by the Joint Collaborative Team on Video Coding (JCT-VC). In the design, tiles are introduced to support high-level parallelism and also to reduce on-chip memory requirements. This paper describes the tile concept and reports results due to the technique.
Arild Fuldseth, Michael Horowitz, Kiran M. Misra, C. Andrew Segall, Minhua Zhou
PCS5
2011 Frame buffer compression for low-power video coding
abstract
In this paper, we propose a novel hybrid frame buffer compression algorithm to reduce the memory bandwidth for lowpower video coding. In our work, we first decompose the full-resolution image into low resolution (LR) and high resolution (HR) components. We then calculate the HR residual by taking the difference between original HR pixel and an estimate derived from surrounding LR pixels. Finally, we use absolute moment block truncation coding to quantize and compress the LR pixel and HR residual data so as to reduce the memory bandwidth. We integrate our approach into the JCT-VC reference software for High Efficiency Video Coding (HEVC). Results show negligible impact on coding efficiency with significant memory bandwidth reduction. Specifically, we observe a bit rate increase of 0.38% and 1% with 21% and 31% memory bandwidth reduction, respectively.
C. Andrew Segall
ICIP2
2011 Periodic entropy coder initialization for wavefront decoding of video bitstream
abstract
In this paper, we outline a coding strategy that initializes the entropy coding engine of a video codec at pre-defined locations within a bit-stream. Coupled with the causal dependencies of state-of-the-art video coding systems, this enables wavefront processing of the entropy decoding and reconstruction process simultaneously. Approaches to wavefront processing have been considered by others, and those methods either address the reconstruction process solely or require transmitting image data in non-raster scan order. Here, our key contribution is that we enable simultaneous entropy/reconstruction wavefront processing while still preserving a raster scan strategy. In this paper, we describe the system, as well as different strategies for initializing the context models. The performance of the proposed methods is evaluated, and the bitrate increase is shown to be nominal.
Kiran M. Misra, Jie Zhao 0007, C. Andrew Segall
ICIP3
2011 Synthesis-Based Texture Video Coding With Side Information
abstract
The addition of a new component to the traditional synthesis-based texture video coding algorithm is investigated in this paper. That is, we add the side information in the form of low-quality video to enhance the texture video synthesis performance with reducing the unpleasant mismatch between analyzed and synthesized regions. As compared with the conventional synthesis algorithm, our algorithm is more flexible since the behavior and quality of the output texture can be adjusted by the amount of the side information, which is determined by the user. To this end, we develop an area-adaptive side information selection scheme that chooses the proper amount of the side information for a given bit budget. Furthermore, we propose a texture decomposition scheme that extracts the non-synthesizable illumination component from the source video for separate coding so as to maximize the synthesis functionality. The superior performance of the proposed texture video synthesis technique is demonstrated by several coding examples.
Byung Tae Oh, Yeping Su, C. Andrew Segall, C.-C. Jay Kuo
IEEE Trans. Circuits Syst. Video Technol.3
2010 Parallel intra prediction for video coding
abstract
In this paper, we propose an intra-prediction system that is both parallel friendly and with high coding efficiency. This is achieved by combining a novel prediction strategy that reduces serial dependencies and a novel, multi-directional and adaptive prediction system. The resulting technique is compared with state-of-the-art ITU-T H.264/MPEG-4 AVC. We observe a 2× and 8× increase in parallelism for 8×8 and 4×4 partitions, respectively, and an average rate increase of less than 0.08% for predictive coding scenarios.
C. Andrew Segall, Jie Zhao 0007, Tomoyuki Yamamoto
PCS1
2009 Video coding for the mobile capture of higher dynamic range image sequences
abstract
This paper is concerned with the problem of capturing higher dynamic range video with a mobile device. We assume the mobile device has a standard (or low) dynamic range image sensor, and that the device is constrained by power and processing capability. To address these issues, we develop a system that captures a video sequence containing time varying exposure settings, encodes this sequence without modification, and then transmits the sequence to a decoder. The bit-stream is constructed so that legacy decoding devices only decode a single exposure setting while advanced devices decode multiple exposure settings and then use the decoded data to reconstruct a higher dynamic range image sequence.
C. Andrew Segall, Jie Zhao 0007, Ron Rubinstein
PCS1
2008 Synthesis-based texture coding for video compression with sideinformation
abstract
This paper presents a synthesis-based texture coding technique that uses low-quality video as side information to control the output texture for video compression. As compared with the current pure synthesis algorithm, the proposed algorithm is generic, in the sense that the behavior and quality of the output texture can be adjusted by the amount of the side information and determined by the user. We develop an area-adaptive side information assignment technique to improve coding efficiency. Additionally, we present a fast-searching algorithm to reduce computational complexity. Simulations demonstrate the performance of the proposed technique.
Byung Tae Oh, Yeping Su, C. Andrew Segall, C.-C. Jay Kuo
ICIP3
2008 Bit stream rewriting for SVC-to-AVC conversion
abstract
The scalable video coding (SVC) extension to the H.264/AVC video coding standard introduces multiple functionalities to the H.264/AVC decoding process. These functionalities include spatial, quality and temporal scalability. In this document, we introduce an additional capability that is supported by the SVC extentions. This feature is commonly referred to as bit-stream rewriting, and it allows a multiple layer, scalable bit-stream to be converted to a single layer, H.264/AVC compliant bit-steam without loss and without reconstruction of image intensity data.
C. Andrew Segall, Jie Zhao 0007
ICIP1
2007 Scalable Coding of High Dynamic Range Video
abstract
A method for coding high dynamic range video sequences is considered. The technique is scalable, in that it facilitates the simultaneous transmission of standard and high dynamic range versions of the sequence in a single bit-stream. Furthermore, the approach is backwards compatible with the existing, state-of-the-art, AVC|H.264 video coding standard. Emphasis is placed on improved coding efficiency as well as managed computational complexity. Results illustrate the efficacy of the approach.
C. Andrew Segall
ICIP (1)1
2007 New Standardized Extensions of MPEG4-AVC/H.264 for Professional-Quality Video Applications
abstract
To support high quality video applications, the Joint Video Team (JVT) has recently added five new profiles, two new supplemental enhancement information (SEI) messages, and two new extended gamut color space indicators to the MPEG4-AVC/H.264 video coding standard. The new profiles include substantial feature enhancements for high-quality video applications, including improved-efficiency 4:4:4 video format coding, improved-efficiency lossless macroblock coding, coding 4:4:4 video pictures using three separately-coded color planes, and support of bit depths up to 14 bits per sample. The new features were developed to support a wide range of applications where high quality video compression is demanded, including professional and semi-professional scenarios in particular. They also anticipate the introduction of higher fidelity displays. In this paper, the new extensions are presented along with quantitaive estimates of the benefits of the new features and a discussion of the target application environments.
Gary J. Sullivan, Haoping Yu, Shun-ichi Sekiguchi, Huifang Sun, Thomas Wedi, Steffen Wittmann, Yung Lyul Lee, C. Andrew Segall, Teruhiko Suzuki
ICIP (1)8
2007 Spatial Scalability Within the H.264/AVC Scalable Video Coding Extension
abstract
A scalable extension to the H.264/AVC video coding standard has been developed within the joint video team (JVT), a joint organization of the ITU-T video coding group (VCEG) and the ISO/IEC moving picture experts group (MPEG). The extension allows multiple resolutions of an image sequence to be contained in a single bit stream. In this paper, we introduce the spatially scalable extension within the resulting scalable video coding standard. The high-level design is described and individual coding tools are explained. Additionally, encoder issues are identified. Finally, the performance of the design is reported.
C. Andrew Segall, Gary J. Sullivan
IEEE Trans. Circuits Syst. Video Technol.1
2006 Resampling for Spatial Scalability
abstract
Resampling is a fundamental issue in the design of a spatially scalable video codec. The resampling procedure is responsible for down-sampling the high-resolution video sequence to generate lower resolution data, as well as upsampling the transmitted lower resolution data to predict the original high-resolution frames. In both cases, the resampling operation must make trade-offs between coding efficiency, image quality and computational complexity. In this paper, we consider the resampling design problem within an optimization framework.
C. Andrew Segall, Aggelos K. Katsaggelos
ICIP1
2004 Improved high-definition video by encoding at an intermediate resolution
abstract
In this paper, we consider the compression of high-definition video sequences for bandwidth sensitive applications. We show that down-sampling the image sequence prior to encoding and then up-sampling the decoded frames increases compression efficiency. This is particularly true at lower bit-rates, as direct encoding of the high-definition sequence requires a large number of blocks to be signaled. We survey previous work that combines a resolution change and compression mechanism. We then illustrate the success of our proposed approach through simulations. Both MPEG-2 and H.264 scenarios are considered. Given the benefits of the approach, we also interpret the results within the context of traditional spatial scalability.
C. Andrew Segall, Michael Elad, Peyman Milanfar, Richard Webb, Chad Fogg
VCIP1
2004 Spatially adaptive high-resolution image reconstruction of DCT-based compressed images
abstract
The problem of recovering a high-resolution image from a sequence of low-resolution DCT-based compressed observations is considered in this paper. The introduction of compression complicates the recovery problem. We analyze the DCT quantization noise and propose to model it in the spatial domain as a colored Gaussian process. This allows us to estimate the quantization noise at low bit-rates without explicit knowledge of the original image frame, and we propose a method that simultaneously estimates the quantization noise along with the high-resolution data. We also incorporate a nonstationary image prior model to address blocking and ringing artifacts while still preserving edges. To facilitate the simultaneous estimate, we employ a regularization functional to determine the regularization parameter without any prior knowledge of the reconstruction procedure. The smoothing functional to be minimized is then formulated to have a global minimizer in spite of its nonlinearity by enforcing convergence and convexity requirements. Experiments illustrate the benefit of the proposed method when compared to traditional high-resolution image reconstruction methods. Quantitative and qualitative comparisons are provided.
Sung Cheol Park, Moon Gi Kang, C. Andrew Segall, Aggelos K. Katsaggelos
IEEE Trans. Image Process.3
2004 Bayesian resolution enhancement of compressed video
abstract
Super-resolution algorithms recover high-frequency information from a sequence of low-resolution observations. In this paper, we consider the impact of video compression on the super-resolution task. Hybrid motion-compensation and transform coding schemes are the focus, as these methods provide observations of the underlying displacement values as well as a variable noise process. We utilize the Bayesian framework to incorporate this information and fuse the super-resolution and post-processing problems. A tractable solution is defined, and relationships between algorithm parameters and information in the compressed bitstream are established. The association between resolution recovery and compression ratio is also explored. Simulations illustrate the performance of the procedure with both synthetic and nonsynthetic sequences.
C. Andrew Segall, Aggelos K. Katsaggelos, Rafael Molina 0001, Javier Mateos
IEEE Trans. Image Process.1
2003 Approaches for the restoration of compressed video
abstract
The restoration of an image sequence that is blurred before compression is considered. This describes many modern imaging systems that filter an image sequence during acquisition and then compress the result. It also describes the common scenario of preprocessing an image sequence with a digital filter prior to compression. No matter the source of degradation though, we seek to recover the high-frequency information without amplifying compression artifacts. The Bayesian framework is employed, and we present recovery algorithms that correspond to two common models for compression noise. Simulations then illustrate the efficacy of both techniques for the restoration of compressed video. Qualitative and quantitative results are presented.
C. Andrew Segall
ICIP (2)1
2002 High-resolution image reconstruction of low-resolution DCT-based compressed images
abstract
The problem of recovering a high-resolution image from a sequence of low-resolution DCT-based compressed images is considered in this paper. The presence of the compression system complicates the recovery problem, as the operation reduces the amount of frequency aliasing in the low-resolution frames and introduces a non-linear quantization process. The effect of the quantization error and resulting inaccurate sub-pixel motion information is modeled as a zero-mean additive correlated Gaussian noise. A regularization functional is introduced not only to reflect the relative amount of registration error in each low-resolution image but also to determine the regularization parameter without any prior knowledge in the reconstruction procedure. The effectiveness of the proposed algorithm is demonstrated experimentally.
Sung Cheol Park, Moon Gi Kang, C. Andrew Segall, Aggelos K. Katsaggelos
ICASSP3
2002 Reconstruction of high-resolution image frames from a sequence of low-resolution and compressed observations
abstract
A framework for recovering high-resolution information from a sequence of sub-sampled and compressed observations is presented. Compression schemes that describe a video sequence through a combination of motion vectors and transform coefficients are the focus (e.g. the MPEG and ITU family of standards), and we consider the influence of both the motion vectors and transform coefficients within the reconstruction algorithm. A Bayesian approach is utilized to incorporate the information, and results show a discemable improvement in resolution, as compared to standard interpolation methods.
C. Andrew Segall, Rafael Molina 0001, Aggelos K. Katsaggelos, Javier Mateos
ICASSP1
2002 Spatially adaptive high-resolution image reconstruction of low-resolution DCT-based compressed images
abstract
The problem of recovering a high-resolution image from a sequence of low-resolution DCT-based compressed images is considered. The presence of the compression system complicates the recovery problem, as the operation reduces the amount of frequency aliasing in the low-resolution frames and introduces a non-linear quantization process, The effect of the quantization error and resulting inaccurate sub-pixel motion information is modeled as a zero-mean additive correlated Gaussian noise. A regularization functional is introduced, not only to reflect the relative amount of registration error in each low-resolution image, but also to determine the regularization parameter without any prior knowledge in the reconstruction procedure. The effectiveness of the proposed algorithm is demonstrated experimentally.
Sung Cheol Park, Moon Gi Kang, C. Andrew Segall, Aggelos K. Katsaggelos
ICIP (2)3
2001 A new constraint for the regularized enhancement of compressed video
abstract
A novel fidelity constraint to the image enhancement problem is presented. With this constraint, we exploit the motion vectors of a compressed video bit-stream. These vectors establish a correspondence between image pixels across a series of frames, and we guarantee that processing the decoded sequence does not violate this correspondence. We develop the constraint within the context of MPEG-2 and incorporate the constraint into a regularized enhancement algorithm. Simulations are then performed. Quantitative and qualitative results illustrate an improvement in visual quality.
C. Andrew Segall, Aggelos K. Katsaggelos
ICASSP1
2001 Bayesian high-resolution reconstruction of low-resolution compressed video
abstract
A method for simultaneously estimating the high-resolution frames and the corresponding motion field from a compressed low-resolution video sequence is presented. The algorithm incorporates knowledge of the spatio-temporal correlation between low and high-resolution images to estimate the original high-resolution sequence from the degraded low-resolution observation. Information from the encoder is also exploited, including the transmitted motion vectors, quantization tables, coding modes and quantizer scale factors. Simulations illustrate an improvement in the peak signal-to-noise ratio when compared with traditional interpolation techniques and are corroborated with visual results.
Rafael Molina 0001, Aggelos K. Katsaggelos, Javier Mateos, C. Andrew Segall
ICIP (2)4
2001 Application of the motion vector constraint to the regularized enhancement of compressed video
abstract
We present a novel fidelity constraint for the image enhancement problem by exploiting the motion vectors of a compressed video bit-stream. These vectors establish a correspondence between image pixels across a series of frames, and our goal is to maintain this relationship during processing. In our past work, we considered algorithms that relied on the sum-of-absolute differences as the match criteria. As we show in this paper, this metric is problematic for the enhancement problem. We then pose the constraint within the context of a sum-of-squared errors criterion for matching. This allows for a more rigorous treatment of the fidelity constraint. Finally, experimental results illustrate the performance of the new constraint.
C. Andrew Segall, Aggelos K. Katsaggelos
ICIP (1)1
2001 A rate-distortion optimal video pre-processing algorithm
abstract
Pre-processing algorithms improve the quality of a compression system by removing unimportant data before encoding. This enhances both the visual quality and coding efficiency of the system. We cast the pre-processing problem in the operational rate-distortion framework. Filtering the displaced frame difference is the focus, and the proposed method couples the choice of the quantization scale to the response of the prefilter. Coding errors are then addressed by penalizing significant differences between coded blocks. Finally, experimental results illustrate the efficacy of the method within the context of an MPEG-2 coding scenario.
C. Andrew Segall, Passant V. Karunaratne, Aggelos K. Katsaggelos
ICIP (1)1
2001 Preprocessing of compressed digital video
C. Andrew Segall, Passant V. Karunaratne, Aggelos K. Katsaggelos
VCIP1
2000 Enhancement of Compressed Video Using Visual Quality Measurements
abstract
The enhancement of compressed video is considered. We present a general algorithm for processing the compressed data, with three variants of the algorithm having practical application. We then consider the algorithm within the context of MPEG-2. Assuming complete knowledge of the compressed bitstream, experiments compare the different realizations of the enhancement algorithm. Our comparisons stress improvements in visual quality, measured by models of the human visual system. Quantitative and qualitative results are provided.
C. Andrew Segall, Aggelos K. Katsaggelos
ICIP1
2000 Restoration of severely blurred high range images using stochastic and deterministic relaxation algorithms in compound Gauss?CMarkov random fields
Rafael Molina 0001, Aggelos K. Katsaggelos, Javier Mateos, Aurora Hermoso, C. Andrew Segall
Pattern Recognit.5
1998 Gradient Independent Translation via Differential Morphology
abstract
A new multi-scale image enhancement mechanism is presented. Derived from the differential representation of the morphological filters, it approximates a median filter and alleviates many of the image blotching and noise preserving characteristics of morphological filtering. In one-dimension, the process is shown to be idempotent and to converge. In two dimensions, experimental results demonstrate convergence and display the ability to remove impulsive noise. Gradient independent translation avoids the two dimensional convergence problems of the median filter and does not involve the expensive rank-ordering of pixel intensities. Describing a scale space with median filter characteristics, it provides a multi-scale analysis method suitable for compression, coding, and feature extraction.
C. Andrew Segall, Scott T. Acton
ICIP (2)1
1997 Morphological Anisotropic Diffusion
abstract
Current formulations of anisotropic diffusion are unable to prevent feature drift and smooth small regions. These deficiencies reduce the effectiveness of the diffusion operation in many image processing tasks, including segmentation, edge detection, compression, and multiscale processing. This paper introduces a morphological diffusion coefficient capable of smoothing small objects while maintaining edge locality. Results are presented that demonstrate its efficacy in edge detection tasks.
C. Andrew Segall, Scott T. Acton
ICIP (3)1