Sriram Sethuraman

dblp:16/3485 · DBLP profile ↗
← Back
26ranked-venue papers
1as first author
16since 2021 · last 2025
0000-0002-2961-1599ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 1 first-author · 16 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Video Quality Assessment for Resolution Cross-Over in Live Sports
abstract
In adaptive bitrate streaming, resolution cross-over refers to the point on the convex hull where the encoding resolution should switch to achieve better quality. Accurate cross-over prediction is crucial for streaming providers to optimize resolution at given bandwidths. Most existing works rely on objective Video Quality Metrics (VQM), particularly VMAF, to determine the resolution cross-over. However, these metrics have limitations in accurately predicting resolution cross-overs. Furthermore, widely used VQMs are often trained on subjective datasets collected using the Absolute Category Rating (ACR) methodologies, which we demonstrate introduces significant uncertainty and errors in resolution cross-over predictions. To address these problems, we first investigate different subjective methodologies and demonstrate that Pairwise Comparison (PC) achieves better cross-over accuracy than ACR. We then propose a novel metric, Resolution Cross-over Quality Loss (RCQL), to measure the quality loss caused by resolution cross-over errors. Furthermore, we collected a new subjective dataset (LSCO) focusing on live streaming scenarios and evaluated widely used VQMs, by benchmarking their resolution cross-over accuracy.
Yixu Chen, Hai Wei, Sriram Sethuraman
ICME4
2024 Encoder-Quantization-Motion-based Video Quality Metrics
abstract
In an adaptive bitrate streaming application, the efficiency of video compression and the encoded video quality depend on both the video codec and the quality metric used to perform encoding optimization. The development of such a quality metric need large scale subjective datasets. In this work we merge several datasets into one to support the creation of a metric tailored for video compression and scaling. We proposed a set of HEVC lightweight features to boost performance of the metrics. Our metrics can be computed from tightly coupled encoding process with 4% compute overhead or from the decoding process in real-time. The proposed method can achieve better correlation than VMAF and P.1204.3. It can extrapolate to different dynamic ranges, and is suitable for real-time video quality metrics delivery in the bitstream. The performance is verified by in-distribution and cross-dataset tests. This work paves the way for adaptive client-side heuristics, real-time segment optimization, dynamic bitrate capping, and quality-dependent post-processing neural network switching, etc.
Yixu Chen, Zaixi Shang, Hai Wei, Sriram Sethuraman
PCS5
2024 HDR-ChipQA: No-reference quality assessment on High Dynamic Range videos
Joshua P. Ebenezer, Zaixi Shang, Hai Wei, Sriram Sethuraman, Alan C. Bovik
Signal Process. Image Commun.5
2024 HDR or SDR? A Subjective and Objective Study of Scaled and Compressed Videos
abstract
We conducted a large-scale study of human perceptual quality judgments of High Dynamic Range (HDR) and Standard Dynamic Range (SDR) videos subjected to scaling and compression levels and viewed on three different display devices. While conventional expectations are that HDR quality is better than SDR quality, we have found subject preference of HDR versus SDR depends heavily on the display device, as well as on resolution scaling and bitrate. To study this question, we collected more than 23,000 quality ratings from 67 volunteers who watched 356 videos on OLED, QLED, and LCD televisions, and among many other findings, observed that HDR videos were often rated as lower quality than SDR videos at lower bitrates, particularly when viewed on LCD and QLED displays. Since it is of interest to be able to measure the quality of videos under these scenarios, e.g. to inform decisions regarding scaling, compression, and SDR vs HDR, we tested several well-known full-reference and no-reference video quality models on the new database. Towards advancing progress on this problem, we also developed a novel no-reference model called HDRPatchMAX, that uses a contrast-based analysis of classical and bit-depth features to predict quality more accurately than existing metrics.
Joshua P. Ebenezer, Zaixi Shang, Yixu Chen, Hai Wei, Sriram Sethuraman, Alan C. Bovik
IEEE Trans. Image Process.6
2024 A Study of Subjective and Objective Quality Assessment of HDR Videos
abstract
As compared to standard dynamic range (SDR) videos, high dynamic range (HDR) content is able to represent and display much wider and more accurate ranges of brightness and color, leading to more engaging and enjoyable visual experiences. HDR also implies increases in data volume, further challenging existing limits on bandwidth consumption and on the quality of delivered content. Perceptual quality models are used to monitor and control the compression of streamed SDR content. A similar strategy should be useful for HDR content, yet there has been limited work on building HDR video quality assessment (VQA) algorithms. One reason for this is a scarcity of high-quality HDR VQA databases representative of contemporary HDR standards. Towards filling this gap, we created the first publicly available HDR VQA database dedicated to HDR10 videos, called the Laboratory for Image and Video Engineering (LIVE) HDR Database. It comprises 310 videos from 31 distinct source sequences processed by ten different compression and resolution combinations, simulating bitrate ladders used by the streaming industry. We used this data to conduct a subjective quality study, gathering more than 20,000 human quality judgments under two different illumination conditions. To demonstrate the usefulness of this new psychometric data resource, we also designed a new framework for creating HDR quality sensitive features, using a nonlinear transform to emphasize distortions occurring in spatial portions of videos that are enhanced by HDR, e.g., having darker blacks and brighter whites. We apply this new method, which we call HDRMAX, to modify the widely-deployed Video Multimethod Assessment Fusion (VMAF) model. We show that VMAF+HDRMAX provides significantly elevated performance on both HDR and SDR videos, exceeding prior state-of-the-art model performance. The database is now accessible at: https://live.ece.utexas.edu/research/LIVEHDR/LIVEHDR_index.html. The model will be made available at a later date at: https://live.ece.utexas.edu//research/Quality/index_algorithms.htm.
Zaixi Shang, Joshua P. Ebenezer, Abhinau Kumar Venkataramanan, Hai Wei, Sriram Sethuraman, Alan C. Bovik
IEEE Trans. Image Process.6
2023 Improving Compression Efficiency using an Encoder-aware Motion Compensated Temporal Filter
abstract
Motion Compensated Temporal Filtering (MCTF) is a pre-processing approach employed prior to video encoding, for improving the compression efficiency. Prior MCTF designs (e.g. [1]) use pre-defined frame-level quantization parameters (QPs) for different slice types and temporal layers, and operate with a fixed Group of Pictures (GOP) structure. However, commercial encoders can adapt GOP structure based upon content characteristics, and can also adapt QPs on a block-basis based upon the frequency of the block being referenced and the spatial complexity of the block, causing prior MCTF to perform sub-optimally with commercial encoders.
Rahul Vanam, Sriram Sethuraman
DCC2
2023 ZREC: Robust Recovery of Mean and Percentile Opinion Scores
abstract
Observer screening and subject opinion score recovery is essential for collecting a reliable QoE database. This paper proposes a new method, ZREC*, which uses Z-scores to estimate subject bias, inconsistency, and content ambiguity. Additionally, we propose Mean Opinion Score (MOS) recovery and Percentile Opinion Score (POS) recovery scheme based on the three estimated parameters. ZREC does not fully reject subjects, rather adjust their coefficients in the MOS/POS recovery, allowing for more efficient use of data collection. The estimated parameters of ZREC are highly correlated with more complex solver-based methods and standards. In addition, ZREC recovers MOS with smaller confidence intervals than the state of the art. Experimental results also demonstrate that using recovered pthPOS as ground truth during training improves the performance of Satisfied User Ratio (SUR) prediction.
Ali Ak, Patrick Le Callet, Sriram Sethuraman, Kumar Rahul
ICIP4
2023 Subjective Test Environments: A Multifaceted Examination of Their Impact on Test Results
abstract
Quality of Experience (QoE) in video streaming scenarios is significantly affected by the viewing environment and display device. Understanding and measuring the impact of these settings on QoE can help develop viewing environment-aware metrics and improve the efficiency of video streaming services. In this ongoing work, we conducted a subjective study in both laboratory and home settings using the same content and design to measure QoE in Degradation Category Rating (DCR). We first analyzed subject inconsistency and confidence intervals of the Mean Opinion Scores (MOS) between the two settings. We then used statistical models such as ANOVA and t-test to analyze the differences in subjective tests on video quality between the two viewing environments. Additionally, we employed the Eliminated-By-Aspects (EBA) model to quantify the influence of different settings on the measured QoE. We conclude with several research questions that could be further explored to better understand the impact of the viewing environment on QoE.
Ali Ak, Charles Dormeval, Patrick Le Callet, Kumar Rahul, Sriram Sethuraman
IMX6
2022 Subjective and Objective Quality Assessment of High-Motion Sports Videos at Low-Bitrates
abstract
Videos often have to be transmitted and stored at low bitrates due to poor network connectivity during adaptive bitrate streaming. Designing optimal bitrate ladders that would select the perceptually-optimized resolution, frame-rate, and compression level for low-bitrate videos for adaptive streaming across the internet is therefore a task of great interest. Towards that end, we conducted the first large-scale study of medium and low-bitrate videos from live sports for two codecs (Elemental AVC and HEVC) and created the Amazon Prime Video Low-Bitrate Sports (APV LBS) dataset. The study involved 94 participants and 742 videos, with more than 23,000 human opinion scores collected in total. We analyzed the data obtained and we also conducted an extensive evaluation of objective Video Quality Assessment (VQA) algorithms and benchmarked their performance, and make recommendations on bitrate ladder design. We're making the metadata and VQA features available at https://github.com/JoshuaEbenezer/lbmfr-public.
Joshua P. Ebenezer, Yixu Chen, Hai Wei, Sriram Sethuraman
ICIP5
2022 Subjective Assessment Of High Dynamic Range Videos Under Different Ambient Conditions
abstract
High Dynamic Range (HDR) videos can represent a much greater range of brightness and color than Standard Dynamic Range (SDR) videos and are rapidly becoming an industry standard. HDR videos have more challenging capture, transmission, and display requirements than legacy SDR videos. With their greater bit depth, advanced electro-optical transfer functions, and wider color gamuts, comes the need for video quality algorithms that are specifically designed to predict the quality of HDR videos. Towards this end, we present the first publicly released large-scale subjective study of HDR videos. We study the effect of distortions such as compression and aliasing on the quality of HDR videos. We also study the effect of ambient illumination on perceptual quality of HDR videos by conducting the study in both a dark lab environment and a brighter living-room environment. A total of 66 subjects participated in the study and more than 20,000 opinion scores were collected, which makes this the largest in-lab study of HDR video quality ever. We anticipate that the dataset will be a valuable resource for researchers to develop better models of perceptual quality for HDR videos.
Zaixi Shang, Joshua P. Ebenezer, Alan C. Bovik, Hai Wei, Sriram Sethuraman
ICIP6
2022 On The Benefit of Parameter-Driven Approaches for the Modeling and the Prediction of Satisfied User Ratio for Compressed Video
abstract
The human eye cannot perceive small pixel changes in images or videos until a certain threshold of distortion. In the context of video compression, Just Noticeable Difference (JND) is the smallest distortion level from which the human eye can perceive the difference between reference video and the distorted/compressed one. Satisfied-User-Ratio (SUR) curve is the complementary cumulative distribution function of the individual JNDs of a viewer group. However, most of the previous works predict each point in SUR curve by using features both from source video and from compressed videos with assumption that the group-based JND annotations follow Gaussian distribution, which is neither practical nor accurate. In this work, we firstly compared various common functions for SUR curve modeling. Afterwards, we proposed a novel parameter-driven method to predict the video-wise SUR from video features. Besides, we compared the prediction results of source-only features based (SRC-based) models and source plus compressed videos features (SRC+PVS-based) models.
Patrick Le Callet, Anne-Flore Perrin, Sriram Sethuraman, Kumar Rahul
ICIP4
2022 Study of the Subjective and Objective Quality of High Motion Live Streaming Videos
abstract
Video livestreaming is gaining prevalence among video streaming service s, especially for the delivery of live, high motion content such as sport ing events. The quality of the se livestreaming videos can be adversely affected by any of a wide variety of events, including capture artifacts, and distortions incurred during coding and transmission. High motion content can cause or exacerbate many kinds of distortion, such as motion blur and stutter. Because of this, the development of objective Video Quality Assessment (VQA) algorithms that can predict the perceptual quality of high motion, live streamed videos is greatly desired. Important resources for developing these algorithms are appropriate databases that exemplify the kinds of live streaming video distortions encountered in practice. Towards making progress in this direction, we built a video quality database specifically designed for live streaming VQA research. The new video database is called the Laboratory for Image and Video Engineering (LIVE) Livestream Database. The LIVE Livestream Database includes 315 videos of 45 source sequences from 33 original contents impaired by 6 types of distortions. We also performed a subjective quality study using the new database, whereby more than 12,000 human opinions were gathered from 40 subjects. We demonstrate the usefulness of the new resource by performing a holistic evaluation of the performance of current state-of-the-art (SOTA) VQA models. We envision that researchers will find the dataset to be useful for the development, testing, and comparison of future VQA models. The LIVE Livestream database is being made publicly available for these purposes at https://live.ece. utexas.edu/research/LIVE_APV_Study/apv_index.html.
Zaixi Shang, Joshua P. Ebenezer, Hai Wei, Sriram Sethuraman, Alan C. Bovik
IEEE Trans. Image Process.5
2021 Detection of Audio-Video Synchronization Errors Via Event Detection
abstract
We present a new method and a large-scale database to detect audio-video synchronization(A/V sync) errors in tennis videos. A deep network is trained to detect the visual signature of the tennis ball being hit by the racquet in the video stream. Another deep network is trained to detect the auditory signature of the same event in the audio stream. During evaluation, the audio stream is searched by the audio network for the audio event of the ball being hit. If the event is found in audio, the neighboring interval in video is searched for the corresponding visual signature. If the event is not found in the video stream but is found in the audio stream, A/V sync error is flagged. We developed a large-scaled database of 504,300 frames from 6 hours of videos of tennis events, simulated A/V sync errors, and found our method achieves high accuracy on the task.
Joshua P. Ebenezer, Hai Wei, Sriram Sethuraman
ICASSP4
2021 Assessment of Subjective and Objective Quality of Live Streaming Sports Videos
abstract
Video live streaming is gaining prevalence among video streaming services, especially for the delivery of popular sporting events. Many objective Video Quality Assessment (VQA) models have been developed to predict the perceptual quality of videos. Appropriate databases that exemplify the distortions encountered in live streaming videos are important to designing and learning objective VQA models. Towards making progress in this direction, we built a video quality database specifically designed for live streaming VQA research. The new video database is called the Laboratory for Image and Video Engineering (LIVE) Live stream Database. The LIVE Livestream Database includes 315 videos of 45 contents impaired by 6 types of distortions. We also performed a subjective quality study using the new database, whereby more than 12,000 human opinions were gathered from 40 subjects. We demonstrate the usefulness of the new resource by performing a holistic evaluation of the performance of current state-of-the-art (SOTA) VQA models. The LIVE Livestream database is being made publicly available for these purposes at https://live.ece.utexas.edu/research/LIVE_APV_Study/apv_index.html.
Zaixi Shang, Joshua P. Ebenezer, Alan C. Bovik, Hai Wei, Sriram Sethuraman
PCS6
2021 Subblock-Based Motion Derivation and Inter Prediction Refinement in the Versatile Video Coding Standard
abstract
Efficient representation and coding of fine-granular motion information is one of the key research areas for exploiting inter-frame correlation in video coding. Representative techniques towards this direction are affine motion compensation (AMC), decoder-side motion vector refinement (DMVR), and subblock-based temporal motion vector prediction (SbTMVP). Fine-granular motion information is derived at subblock level for all the three coding tools. In addition, the obtained inter prediction can be further refined by two optical flow-based coding tools, the bi-directional optical flow (BDOF) for bi-directional inter prediction and the prediction refinement with optical flow (PROF) exclusively used in combination with AMC. The aforementioned five coding tools have been extensively studied and finally adopted in the Versatile Video Coding (VVC) standard. This paper presents technical details of each tool and highlights the design elements with the consideration of typical hardware implementations. Following the common test conditions defined by Joint Video Experts Team (JVET) for the development of VVC, 5.7% bitrate reduction on average is achieved by the five tools. For test sequences characterized by large and complex motion, up to 13.4% bitrate reduction is observed. Additionally, visual quality improvement is demonstrated and analyzed.
Haitao Yang 0001, Huanbang Chen, Jianle Chen, Semih Esenlik, Sriram Sethuraman, Xiaoyu Xiu, Elena Alshina, Jiancong Luo
IEEE Trans. Circuits Syst. Video Technol.5
2021 ChipQA: No-Reference Video Quality Prediction via Space-Time Chips
abstract
We propose a new model for no-reference video quality assessment (VQA). Our approach uses a new idea of highly-localized space-time (ST) slices called Space-Time Chips (ST Chips). ST Chips are localized cuts of video data along directions that implicitly capture motion. We use perceptually-motivated bandpass and normalization models to first process the video data, and then select oriented ST Chips based on how closely they fit parametric models of natural video statistics. We show that the parameters that describe these statistics can be used to reliably predict the quality of videos, without the need for a reference video. The proposed method implicitly models ST video naturalness, and deviations from naturalness. We train and test our model on several large VQA databases, and show that our model achieves state-of-the-art performance at reduced cost, without requiring motion computation.
Joshua P. Ebenezer, Zaixi Shang, Hai Wei, Sriram Sethuraman, Alan C. Bovik
IEEE Trans. Image Process.5
2008 A low-complexity, motion-robust, spatio-temporally adaptive video de-noiser with in-loop noise estimation
abstract
Noise in video influences the bit-rate and visual quality of video encoders and can significantly alter the effectiveness of video processing algorithms. Recent advances offer high quality de-noising at a fairly high computational complexity by increasing the spatio-temporal support and evaluating intelligent weights for combining these samples to remove noise while preserving the signal. The lower complexity methods, typically, either over-blur the video or introduce motion artifacts and temporal flicker. By reusing the motion vectors generated by a video encoder, a low incremental complexity de-noiser is proposed in this paper that is capable of achieving a high level of noise reduction and signal preservation with a reduced spatio-temporal support. In addition, the approach lends itself to dynamically estimating the noise variance used for controlling the level of filtering. The proposed approach performs on par with the spatial non-local means de-noising algorithm for stationary background sequences and can be improved for motion sequences with motion-compensated temporal filtering.
Sriram Sethuraman
ICIP2
2008 Dynamic frame-rate selection for live LBR video encoders using trial frames
abstract
Low bit-rate live encoding and streaming of video over wired or wireless channels, as utilized in place-shifting applications, need to adapt the frame rate and bit-rate according to content and channel conditions. Typical rate control schemes work with a given target frame rate and can result in irregular frame skips that are induced by buffer conditions and quantizer saturation. Prior methods proposed to address the dynamic frame-rate selection (DFS) problem work by projecting the impact of a new frame-rate from the operating frame-rate and can have slower convergence to the required frame-rate. In this paper, we propose a scheme that utilizes a given computational capacity (typical in embedded appliances) to actually assess the impact of the new frame-rate by coding at the corresponding new frame-skip (referred to as trial frames) before committing to it. The proposed scheme offers quick content-adaptive convergence to the target frame-rate and maintains a motion-adaptive frame skipping pattern to provide a subjectively pleasing spatio-temporal quality trade-off at the operating bit-rate. In addition, the scheme can work with most existing rate control methods that operate at a supplied frame-rate and is also suitable for hooking to a user controlled quality dial.
Arunoday Thammineni, Arvind Raman, Sarat Chandra Vadapalli, Sriram Sethuraman
ICME4
2008 Low-complexity frame-level joint source-channel distortion optimal, adaptive intra refresh
abstract
Error resilient, low latency video coding for interactive video applications requires progressive intra coding of macroblocks (MBs) to contain the error propagation. Both refresh MB selection (RMS) and refresh rate selection (RRS) impact the subjective video quality in the presence of packet losses. Joint source-channel rate distortion optimization methods attempt to find the best trade-off between compression efficiency and end-to-end distortion at an MB-level and are typically computationally expensive in addition to not being optimal at a picture level. While probabilistic error propagation tracking is used for refresh MB selection in previous work, these picture-level optimal RRS methods model source-channel distortion by mimicking the effect of periodic intra frame coding which does not match well with content adaptive refresh MB selection. In this paper, we propose a frame-level approach to RRS that aligns the joint source-channel rate-distortion trade-off modeling with an enhanced RMS process to achieve an optimal end-to-end distortion that is content, bit-rate and channel adaptive. For typical videoconferencing content, the proposed approach is quite low in complexity, works on par with off-line multi-pass identification of the optimal fixed refresh rate and is quite competitive when compared to the H.264 joint modelpsilas lossy rate-distortion optimization technique.
Sarat Chandra Vadapalli, Biswadeep Sengupta, Sriram Sethuraman
MMSP3
2008 Making Video Quality Assessment Models Robust to Bit Depth
abstract
We introduce a novel feature set, which we call HDRMAX features, that when included into Video Quality Assessment (VQA) algorithms designed for Standard Dynamic Range (SDR) videos, sensitizes them to distortions of High Dynamic Range (HDR) videos that are inadequately accounted for by these algorithms. While these features are not specific to HDR, and also augment the equality prediction performances of VQA models on SDR content, they are especially effective on HDR. HDRMAX features modify powerful priors drawn from Natural Video Statistics (NVS) models by enhancing their measurability where they visually impact the brightest and darkest local portions of videos, thereby capturing distortions that are often poorly accounted for by existing VQA models. As a demonstration of the efficacy of our approach, we show that, while current state-of-the-art VQA models perform poorly on 10-bit HDR databases, their performances are greatly improved by the inclusion of HDRMAX features when tested on HDR and 10-bit distorted videos.
Joshua P. Ebenezer, Zaixi Shang, Hai Wei, Sriram Sethuraman, Alan C. Bovik
IEEE Signal Process. Lett.5
2007 Efficient Alternative to Intra Refresh using Reliable Reference Frames
abstract
Motion compensated prediction of video results in decoder-side error propagation (EP) in the presence of channel losses. Progressive intra-refresh methods control the EP, but lead to a drop in coding efficiency (CE). Error-resilient video coding schemes using multiple reference frames have been proposed earlier to achieve a better CE. Some schemes (e.g. ACK-based) tend to be conservative at the expense of CE. Methods using probabilistic models and motion tracking to estimate the decoder state offer better CE, but do not take care of the time taken to recover from losses and the EP due to the use of unreliable reference frames. In this paper, we propose an alternative to intra-refresh that controls the EP and overcomes the CE loss through the use of reliable (i.e. near drift-free) reference frames to code the refresh macroblocks. The proposed scheme has been validated through integration into a ITU-T H.264 encoder and has been shown to give a CE gain of up to 1.5dB. Further, it is of low-complexity and suited for implementation on embedded platforms.
Sarat Chandra Vadapalli, Harish Shetiya, Sriram Sethuraman
ICME3
2004 Low-cost wireless projector interface device using TI TMS320DM270
abstract
Personal computers today have 802.11 (b/a/g) -based wireless connectivity. By connecting laptops to a projector wirelessly, multiple persons could connect to the projector and present the content on their laptops. This typically requires another personal computer at the projector end that is wirelessly connected to the laptop and which receives, decodes, and renders a compressed version of what was rendered on the display of the laptop. On the laptop, the screen buffer is captured and the captured frames are compressed. The encoded stream is streamed over the WLAN to a streaming client. The article presents a low cost interface device for the projector system based on the TI TMS320DM270 image/video processor.
Arvind Raman, Mini Jain, T. C. Rajendra, S. Satheesh, Sriram Sethuraman, Vikal Kumar Jain, Vinayak Prasanna Das
ICME5
2004 Depth map compression for real-time view-based rendering
Bing-Bing Chai, Sriram Sethuraman, Harpreet Sawhney, Paul Hatrack
Pattern Recognit. Lett.2
2002 Mesh-based depth map compression and transmission for real-time view-based rendering
abstract
Enabling telepresence using depth-based new view rendering requires the compression and transmission of dynamic depth maps and video from multiple cameras. The telepresence application places additional requirements on the compressed representation, such as preservation of depth discontinuities, low complexity decoding, and amenability to real-time rendering using graphics cards. We propose a simplified triangular mesh representation that can be encoded efficiently. The mesh geometry is encoded using a binary triangle tree structure and the depths at the tree nodes are differentially coded. By matching the tree traversal to the mesh rendering order, both depth map decoding and triangle strip generation for efficient rendering are achieved simultaneously. The proposed scheme naturally lends itself to coding segmented foreground layers. Our experiments show a significant improvement in rendering speeds at similar compression rates using the new mesh based depth map representation when compared to independent representations for compression and rendering.
Bing-Bing Chai, Sriram Sethuraman, Paul Hatrack
ICIP (2)2
2001 Compression and transmission of depth maps for image-based rendering
abstract
We consider applications using depth-based image-based rendering (IBR), where the synthesis of arbitrary views occur at a remote location, necessitating the compression and transmission of depth maps. Traditional image compression has been designed to provide maximum perceived visual quality, and a direct application is sub-optimal for depth-map compression, since depth-maps are not directly viewed. In other words, the sensitivity of the rendering error depends on the image content as well as on the depth map, we propose two improvements to take this into account. Firstly, we consider region-of-interest (ROI) coding, where we identify those regions of the image where accurate depth is most crucial. Secondly, we reshape the dynamic range of the depth map. Our experiments show a significant improvement in coding gain (1.1 dB) and rendering quality when we integrated these two improvements into a standard JPEG-2000 coder.
Ravi Krishnamurthy, Bing-Bing Chai, Sriram Sethuraman
ICIP (3)4
1994 A Multiresolution Framework for Stereoscopic Image Sequence Compression
abstract
Stereoscopic sequence compression typically involves the exploitation of the spatial redundancy between the left and right streams to achieve higher compressions than are possible with the independent compression of the two streams. In this paper the psychophysical property of the human visual system, that only one high resolution image in a stereo image pair is sufficient for satisfactory depth perception, has been used to further reduce the bit rates. Thus, one of the streams is independently coded along the lines of the MPEG standards, while the other stream is estimated at a lower resolution from this stream. A multiresolution framework has been adopted to facilitate such an estimation of motion and disparity vectors at different resolutions. Experimental results on typical sequences indicate that the additional stream can be compressed to about one-fifth of a highly compressed independently coded stream, without any significant loss in depth perception or perceived image quality.>
Sriram Sethuraman, Mel W. Siegel, Angel G. Jordan
ICIP (2)1