EDBT 2026 Demo / reviewers in the wild / expert
Bharath Vishwanath
dblp:201/7092
· DBLP profile ↗
22ranked-venue papers
18as first author
11since 2021 · last 2025
0000-0002-3098-9097ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 18 first-author · 10 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cross-Component Residual Prediction for Geometry-Based Point Cloud CompressionabstractPoint cloud compression is pivotal for the success of immersive multimedia applications. For attribute compression in geometry-based point cloud compression (G-PCC), Region Adaptive Hierarchical Transform (RAHT) is the preferred coding method. Inspired by the significant impact of cross-component prediction in traditional image and video coding, we investigate and present our pioneering work on cross-component residual prediction for RAHT in G-PCC. The method builds on the core observation that cross-component correlations are observed locally in some regions in some sequences. Accordingly, the prediction is employed for last few layers of RAHT which capture local characteristics. We employ a simple linear model, that predicts chroma residues from reconstructed luma residue. The prediction coefficients are learnt on the fly from reconstructed residues of the neighbors. The method gives 1% luma coding gain and around 2-3% chroma coding gain with negligible increase in complexity. The method is adopted to the Geometric Solid Test Model (GeS-TM v7.0), a dedicated codec being developed for solid point clouds. Bharath Vishwanath, Yingzhan Xu, Kai Zhang 0007, Li Zhang 0136 |
ICASSP | 1 |
| 2025 | Rate-Distortion Optimized Chroma Quantization for Point Cloud CompressionabstractPoint cloud compression is pivotal for the success of immersive multimedia applications. For attribute compression in the geometry-based point cloud compression (G-PCC), Region Adaptive Hierarchical Transform (RAHT) is the preferred coding method. G-PCC performs rate-distortion optimized quantization of luma and chroma residue where they are jointly quantized to zero if deemed to be R-D optimal. However, it is often beneficial to zero out the chroma residue while retaining the luma residue, since chroma exhibits less variations. To address this, we propose rate-distortion (R-D) optimized chroma quantization. Each chroma residue sample is decided to be quantized to zero based on the R-D cost. For accurate R-D cost evaluation, we propose a method to estimate the Lagrange multiplier λ on the fly and further scale it according to the RAHT layers, achieving sequence and layer-wise adaptivity. The method gives 1% effective luma coding gain with negligible increase in complexity. The method has been adopted to the next version of Geometric Solid Test Model (GeS-TM v8.0), a dedicated codec being developed for solid point clouds. Bharath Vishwanath, Yingzhan Xu, Kai Zhang 0007 |
ICIP | 1 |
| 2025 | Advances in Predictive RAHT for Geometric Point Cloud CompressionabstractPoint cloud compression is critical for the success of immersive multimedia applications. For attribute compression in geometric point cloud compression (G-PCC), Region Adaptive Hierarchical Transform (RAHT) is the preferred coding method. This paper presents several advances to predictive coding with RAHT: 1) Sample Domain Prediction: Prediction in RAHT is done in transform domain. This introduces undesirable distortion to the prediction signal because of fixed-point computations and leads to increased decoding complexity. We address this by naturally applying prediction in sample domain. The method opens door to skip the transform stage altogether when all residues are quantized to zero, leading to a significantly light decoder. 2) Reference Node Resampling: Inter-prediction signal derived in RAHT could have a different occupancy and weight distribution compared to the current block, causing a mismatch. To address this, we resample the reference node and align the occupancy and weight distribution. 3) Temporal Filtering: During inter-prediction, the reference node is simply copied as the prediction signal. This assumes a correlation coefficient of unity, which is barely true. We introduce a temporal filtering mechanism conditioned on the sub-band, that emulates a low-pass filtering and achieves improved prediction. 4) Inter-Eligibility: During AC inter-prediction, both encoder and decoder have access to the DC of the current and the reference nodes. We use this information to derive an inter-eligibility criterion. Experimental results show considerable gains and reduced complexity that demonstrate the utility of the proposed methods. All the presented methods have been adopted to the second version of G-PCC. Bharath Vishwanath, Kai Zhang 0007, Li Zhang 0006 |
IEEE Trans. Image Process. | 1 |
| 2024 | Sample Domain Prediction and Transform Skip for Region Adaptive Hierarchical Transform in Geometric Point Cloud CompressionabstractPoint cloud compression is critical for the success of immersive multimedia applications. For attribute compression in geometric point cloud compression (G-PCC), Region Adaptive Hierarchical Transform (RAHT) is the preferred coding method. Although RAHT was initially introduced as a pure transform coding tool, recent advancements have introduced intra and inter prediction for RAHT. However, these methods perform prediction in transform domain which is sub-optimal since: ${i}$) fixed-point RAHT introduces distortion to the prediction signal and $i {i}$) transforming prediction signal leads to additional decoding complexity. To address this, we propose to perform prediction in sample domain, thereby retaining crisp prediction signal and alleviating decoder of unnecessary computations. Performing prediction in sample domain opens door to completely skip the transform stage at the decoder when all the residue of a block are quantized to zero, leading to further complexity reduction. The proposed methods achieve an average chroma coding gain of around $1 \%$ and reduces the overall decoding complexity by $3-5 \%$. The method is adopted to the next version of Geometric Solid Test Model (GeS-TM v5.0) and is being evaluated on G-PCC test model TMC13v25. Bharath Vishwanath, Yingzhan Xu, Kai Zhang 0007, Li Zhang 0136 |
ICIP | 1 |
| 2024 | Improved Geometry Coding for Spinning LiDAR Point Cloud CompressionabstractPoint cloud compression has emerged as a hot research topic in recent years. Due to applications such as autonomous driving, LiDAR point cloud compression is an important research aspect of this field. Moving Picture Experts Group (MPEG) is developing a standard called Geometry-based Point Cloud Compression (G-PCC) to meet the compression requirements of point clouds from different collection devices including LiDAR. In current G-PCC, the prior information of spinning LiDAR is not fully utilized in octree geometry coding. In this paper, we address this issue and effectively account for the prior information of spinning LiDAR to improve the compression efficiency of octree geometry coding. Specifically, the angle information provided by capture laser scanner is utilized for Inferred Direct Coding Mode (IDCM) eligibility criterion and z coordinate compensation of the reconstructed point cloud. Experimental results demonstrate that the proposed method achieves 6.7% and 16.4% average coding gain under D1 and D2 quality metrics, respectively, with a negligible increase in complexity. The major part of the proposed method has been adopted in G-PCC. Yingzhan Xu, Bharath Vishwanath, Kai Zhang 0007, Li Zhang 0136 |
ISCAS | 3 |
| 2024 | A Discrete-Mapping-Based Cross-Component Prediction Paradigm for Screen Content CodingabstractCross-component prediction is an important intra-prediction tool in the modern video coders. Existing prediction methods to exploit cross-component correlation include cross-component linear model and its extension of multi-model linear model. These models are designed for camera captured content. For screen content coding, where videos exhibit different signal characteristics, a cross-component prediction model tailored to their characteristics is desirable. As a pioneering work, we propose a discrete-mapping based cross-component prediction model for screen content coding. Our model relies on the core observation that, screen content videos typically comprise of regions with a few distinct colors and luma value (almost always) uniquely conveys chroma value. Based on this, the proposed method learns a discrete-mapping function from available reconstructed luma-chroma pairs and uses this function to derive chroma prediction from the co-located luma samples. To achieve higher accuracy, a multi-filter approach is employed to derive co-located luma values. The proposed method achieves 2.61%, 3.51% and 3.92% Y, U and V bit-rate savings respectively over Enhanced Compression Model (ECM) 4.0, with negligible complexity, for text and graphics media under all-intra configuration. Bharath Vishwanath, Kai Zhang 0007, Li Zhang 0006 |
IEEE Trans. Image Process. | 1 |
| 2023 | Temporal Filtering for Region Adaptive Hierarchical Transform in Geometric Point Cloud CompressionabstractDynamic point cloud compression is critical for the success of immersive multimedia applications and autonomous driving. For attribute compression in geometric point cloud compression (G-PCC), Region Adaptive Hierarchical Transform (RAHT) is the preferred coding method. Recently, inter-prediction was introduced for RAHT in G-PCC. In inter-RAHT, the transform coefficients of the top layers (low frequency coefficients) are predicted by a simple copy of the transform coefficients from the reference frame. Such a prediction assumes reference and current RAHT layers are perfectly correlated, which is barely true. To address this, the paper introduces a temporal filtering mechanism for RAHT. Specifically, we scale the reference layer by a filtering coefficient. Different filtering coefficients are designed for different RAHT layers, thereby emulating a low-pass filter. Experiments show an average 2% bit-rate savings over G-PCC test model tmc13-v22 with negligible increase in complexity. The method has been adopted to the second version of G-PCC. Bharath Vishwanath, Yingzhan Xu, Kai Zhang 0007, Li Zhang 0136 |
VCIP | 1 |
| 2022 | Joint Asymptotic Closed-Loop Design of Secondary Transform and Scan Order for Inter Coding in AV1abstractMost video coding systems employ separable transforms due to their low computational complexity and storage requirements, and despite their sub-optimal decorrelating capabilities. To achieve better decorrelation, recent codecs further apply a non-separable secondary transform to low-frequency primary transform coefficients of the intra-prediction residual. This paper focuses on effective design of non-separable secondary transforms for the inter-prediction residual. As the combination of primary and secondary transforms yields a set of ultimate transform coefficients for which the default zig-zag scan order is sub-optimal, we complement the secondary transform design with the design of corresponding coefficient scanning order modes that facilitate effective entropy coding. A critical stability challenge in this joint design, due to error propagation through the codec's prediction loop, is circumvented by leveraging the asymptotic closed loop (ACL) design paradigm. ACL operates in open-loop in each iteration to ensure stability, but in a manner that ultimately converges to parameter optimization for closed-loop operation. Experimental results show an average gains in BD rate of 0.76% for CIF and 0.41% for HD sequences, over the AV1 codec (libaom). Kruthika Koratti Sivakumar, Bharath Vishwanath, Kenneth Rose |
MMSP | 2 |
| 2022 | A Cross-Component Prediction Model For Screen Content CodingabstractCurrent video coding schemes such as VVC and ECM employ separate palette coding for luma and chroma components under dual-tree structure, ignoring cross-component correlations. Although there are linear and multi-modal linear models to capture cross-component correlations, such models are not tailored for screen content sequences. To address this, we propose a novel cross-component prediction model for screen content sequences. The proposed method builds on the core observation that, regions of screen content sequences comprise of few distinct colors and luma value (almost always) uniquely conveys chroma values. In the light of this observation, the proposed method derives chroma prediction based on a discrete mapping function between luma and chroma values. Specifically, the method simply remembers the reconstructed luma values and their corresponding chroma values in a look-up table and employs this look-up table for cross-component prediction for the current chroma block. To achieve higher accuracy, a multi-filter approach is employed to derive co-located luma values. For an example configuration, the proposed method achieves 1.37%, 1.08% and 1.68% Y, U and V bit-rate savings respectively over ECM 3.1, for text and graphics media under all-intra configuration, demonstrating its efficacy. Bharath Vishwanath, Kai Zhang 0007, Li Zhang 0006 |
PCS | 1 |
| 2022 | Effective Prediction Modes Design for Adaptive Compression With Application in Video CodingabstractAdaptive prediction is an important tool for efficient compression of non-stationary signals. A common approach to achieve adaptivity is to switch between a set of prediction modes, designed to capture variations in signal statistics. The design poses several challenges including: i) catastrophic instability due to statistical mismatch driven by propagation through the prediction loop, and ii) severe non-convexity of the cost surface that is often riddled with poor local minima. Motivated by these challenges, this paper presents a near-optimal method for designing prediction modes for adaptive compression. The proposed method builds on a stable, open-loop platform, but with a subterfuge that ensures that the design is asymptotically optimized for closed-loop operation. The non-convexity is handled by deterministic annealing, a powerful optimization tool to avoid poor local minima. To demonstrate the impact of the proposed approach on practical applications, we consider adaptive, transform-domain predictor design for enhancing standard video coding. Experimental results validate the benefits of the proposed design in terms of significant performance gains for both predictive compression systems in general and video coding in particular. Bharath Vishwanath, Tejaswi Nanjundaswamy, Kenneth Rose |
IEEE Trans. Image Process. | 1 |
| 2022 | A Geodesic Translation Model for Spherical Video CompressionabstractSpherical video coding is critical to the success of many virtual reality and related applications. This paper focuses on an important class of spherical videos whose dynamics involve camera motion. A common approach to spherical video coding is to project from the sphere onto a plane (or planes), where a standard video coder is applied. The projection induces warping resulting in complex non-linear motion in the projected domain that severely comprises the performance of motion models in standard coders. To overcome this shortcoming, we propose a new motion model that captures the motion field on the sphere, and capitalizes on insights into the perceived motion on the sphere due to camera translation. Specifically, surrounding static points are perceived as moving along their respective geodesics, which all intersect at the points where the camera velocity vector intersects the sphere. We analyze the rate of translation along geodesics and its dependence on the elevation of a pixel on the sphere with respect to the camera velocity vector. The analysis leads to a motion vector modulation scheme that perfectly captures the perceived motion of each pixel. Complementary to the new motion model, we propose a search grid tailored to capture expected geodesic motion on the sphere for effective motion estimation. The proposed method yields significant bit-rate savings over employing standard HEVC after projection, which validates its efficacy. Bharath Vishwanath, Tejaswi Nanjundaswamy, Kenneth Rose |
IEEE Trans. Image Process. | 1 |
| 2020 | Spherical Video Coding with Geometry and Region Adaptive Transform Domain Temporal PredictionabstractMany virtual and augmented reality applications depend critically on efficient compression of spherical videos. Current approaches apply a projection geometry to map a spherical video onto the plane(s), wherein a standard codec can be used for compression. Video coders employ simple pixel copying from reference frames for inter-prediction, which ignores underlying spatial correlations, and is hence suboptimal. A novel paradigm of transform domain temporal prediction (TDTP) was developed previously in our lab to effectively overcome this suboptimality of standard video coding. This paper is motivated by the observation that projected spherical videos exhibit significantly more statistical variation due to i) the choice of projection geometry and ii) position of the block on the sphere, which reflect variations in sampling density and various statistical features. To account for such variations, we propose geometry and region adaptive TDTP that is tailored to spherical videos. For a given geometry, the sphere is divided into regions, according to expected signal statistics, and prediction filters are designed for each region. Experimental results show significant performance gains as evidence for the efficacy of TDTP in spherical video coding. Bharath Vishwanath, Kenneth Rose |
ICASSP | 1 |
| 2020 | Asymptotic Closed-Loop Design Of Transform Modes For The Inter-Prediction Residual In Video CodingabstractTransform coding is a key component of video coders, tasked with spatial decorrelation of the prediction residual. There is growing interest in adapting the transform to local statistics of the inter-prediction residual, going beyond a few standard trigonometric transforms. However, the joint design of multiple transform modes is highly challenging due to critical stability problems inherent to feedback through the codec's prediction loop, wherein training updates inadvertently impact the signal statistics the transform ultimately operates on, and are often counter-productive (and sometimes catastrophic). It is the premise of this work that a truly effective switched transform design procedure must account for and circumvent this shortcoming. We introduce a data-driven approach to design optimal transform modes for adaptive switching by the encoder. Most importantly, to overcome the critical stability issues, the approach is derived within an asymptotic closed loop (ACL) design framework, wherein each iteration operates in an effective open loop, and is thus inherently stable, but with a subterfuge that ensures that, asymptotically, the design approaches closed loop operation, as required for the ultimate coder operation. Experimental results demonstrate the efficacy of the proposed optimization paradigm which yields significant performance gains over the state-of-the-art. Bharath Vishwanath, Shunyao Li, Kenneth Rose |
ICIP | 1 |
| 2020 | Geodesic Disparity Compensation for Inter-View Prediction in VR180abstractThe VR180 format is gaining considerable traction among the various promising immersive multimedia formats that will arguably dominate future multimedia consumption applications. VR180 enables stereo viewing of a hemisphere about the user. The increased field of view and the stereo setting result in extensive volumes of data that strongly motivate the pursuit of novel efficient compression tools tailored to this format. This paper's focus is on the critical inter-view prediction module that exploits correlations between camera views. Existing approaches mainly consist of projection to a plane where traditional multi-view coders are applied, and disparity compensation employs simple block translation in the plane. However, warping due to the projection renders such compensation highly suboptimal. The proposed approach circumvents this shortcoming by performing geodesic disparity compensation on the sphere. It leverages the observation that, as an observer moves from one view point to the other, all points on surrounding objects are perceived to move along respective geodesics on the sphere, which all intersect at the two points where the axis connecting the two view points pierces the sphere. Thus, the proposed method performs inter-view prediction on the sphere by moving pixels along their predefined respective geodesics, and accurately captures the perceived deformations. Experimental results show significant bitrate savings and evidence the efficacy of the proposed approach. Kruthika Koratti Sivakumar, Bharath Vishwanath, Kenneth Rose |
VCIP | 2 |
| 2019 | Deterministic Annealing Based Transform Domain Temporal Predictor Design for Adaptive Video CodingabstractCurrent video coders employ motion compensated pixel-to-pixel prediction, which largely ignores significant spatial correlations and the fact that true temporal correlations vary with spatial frequency. Earlier work from our lab proposed to first spatially decorrelate the block of pixels by performing temporal prediction in the transform domain, and to effectively account for both spatial and temporal correlations. To adapt to variations in video signal statistics, the encoder switches between a set of appropriately designed prediction modes.This setting critically depends on efficient offline learning of transform domain temporal prediction modes. Significant challenges include: i) issues of instability and mismatched statistics inherent to closed loop design; and ii) severe non-convexity of the cost function trapping the system in poor local minima. Statistics mismatch is tackled by an appropriate paradigm for system design in a stable open loop fashion, but which asymptotically mimics closed loop operation. The non-convexity is handled by deterministic annealing, a powerful non-convex optimization tool whose probabilistic formulation allows for direct optimization of the cost function with respect to the discrete set of prediction modes, and whose annealing schedule avoids poor local minima. Experimental results validate the method's efficacy. Bharath Vishwanath, Tejaswi Nanjundaswamy, Kenneth Rose |
DCC | 1 |
| 2019 | A Deterministic Annealing Approach to Switched Predictor Design for Adaptive Compression SystemsabstractAdaptive prediction is important in the compression of non-stationary signals, and a common remedy is to switch between appropriately designed prediction modes. This paper presents a near optimal procedure to design prediction modes for an adaptive compression system. The main challenges include: instability and mismatched statistics during closed loop design; and the severe non-convexity of the cost function trapping the system in poor local minima. The statistical mismatch is circumvented through a largely open loop (hence stable) design that is devised to asymptotically optimize the prediction modes for closed loop operation. The non-convexity of the cost function is handled by the deterministic annealing paradigm, a powerful non-convex optimization framework devised to avoid poor local minima. Experimental results provide substantial gains validating the efficacy of the proposed design technique. Bharath Vishwanath, Tejaswi Nanjundaswamy, Kenneth Rose |
ICASSP | 1 |
| 2019 | Spherical Video Coding With Motion Vector Modulation to Account For Camera MotionabstractEmerging immersive multimedia applications critically depend on efficient compression of spherical (360-degree) videos. Current approaches project spherical video onto planes for coding with standard codecs, without accounting for the properties of spherical video, a severe sub-optimality that motivates this work. A common type of spherical video is dominated by camera translation. We recently proposed a powerful motion compensation technique for such videos which builds on the observation that, with camera translation, stationary points are perceived as moving along geodesics that meet at the point where the camera translation vector intersects the sphere. However, the approach follows standard coding procedures and translates all pixels in a block by the same amount on their respective geodesics, which is sub-optimal. This paper analyzes the appropriate rate of translation along geodesics and its dependence on the elevation of a pixel on the sphere with respect to the camera velocity pole. The analysis leads to a new approach that modulates the effective motion vectors within a block such that they perfectly capture the perceived individual motion of each pixel. Consistent gains in the experiments provide evidence for the efficacy of the proposed approach. Bharath Vishwanath, Kenneth Rose |
VCIP | 1 |
| 2018 | Motion Compensated Prediction for Translational Camera Motion in Spherical Video CodingabstractSpherical video is the key driving factor for the growth of virtual reality and augmented reality applications, as it offers truly immersive experience by capturing the entire 3D surroundings. However, it represents an enormous amount of data for storage/transmission and success of all related applications is critically dependent on efficient compression. A frequently encountered type of content in this video format is due to translational motion of the camera (e.g., a camera mounted on a moving vehicle). Existing approaches simply project this video onto a plane and use block based translational motion model for capturing the motion of the objects between the frames. This ad-hoc simplified approach completely ignores the complex deformities of objects caused due to the combined effect of the moving camera and projection onto a plane, rendering it significantly suboptimal. In this paper, we provide an efficient solution tailored to this problem. Specifically, we propose to perform motion compensated prediction by translating pixels along their geodesics, which intersect at the poles corresponding to the camera velocity vector. This setup not only captures the surrounding objects' motion exactly along the geodesics of the sphere, but also accurately accounts for the deformations caused due to projection on the sphere. Experimental results demonstrate that the proposed framework achieves very significant gains over existing motion models. Bharath Vishwanath, Tejaswi Nanjundaswamy, Kenneth Rose |
MMSP | 1 |
| 2018 | Rotational Motion Compensated Prediction in HEVC Based Omnidirectional Video CodingabstractSpherical video is becoming prevalent in virtual and augmented reality applications. With the increased field of view, spherical video needs enormous amounts of data, obviously demanding efficient compression. Existing approaches simply project the spherical content onto a plane to facilitate the use of standard video coders. Earlier work at UCSB was motivated by the realization that existing approaches are suboptimal due to warping introduced by the projection, yielding complex non-linear motion that is not captured by the simple translational motion model employed in standard coders. Moreover, motion vectors in the projected domain do not offer a physically meaningful model. The proposed remedy was to capture the motion directly on the sphere with a rotational motion model, in terms of sphere rotations along geodesics. The rotational motion model preserves the shape and size of objects on the sphere. This paper implements and tests the main ideas from the previous work [1] in the context of a full-fledged, unconstrained coder including, in particular, bi-prediction, multiple reference frames and motion vector refinement. Experimental results provide evidence for considerable gains over HEVC. Bharath Vishwanath, Kenneth Rose, Yuwen He, Yan Ye 0003 |
PCS | 1 |
| 2017 | Deterministic annealing based design of error resilient predictive compression systemsabstractThis paper considers near optimal design of predictive compression system that accounts for packet loss over unreliable networks. Major challenges to address include, propagation of errors due to packet loss through the prediction loop, mismatch between statistics used for design and during operation, and above all a cost function that is fraught with poor local minima. Accurately estimating and minimizing the end-to-end distortion (EED), in combination with asymptotic closed-loop (ACL) design that employs open-loop iterations, but mimics closed-loop operation on convergence, was proposed to address the first two challenges. However the severe non-convexity of the cost function, especially due to the piece-wise linear nature of the quantizer function, makes this a particularly challenging optimization problem. We propose to tackle this via a new design approach in the deterministic annealing framework to avoid poor local minima, coupled with the ACL approach to minimize EED estimate. This effectively addresses all the major challenges, and leads to a near optimal design of error-resilient predictive compression system. Substantial performance improvement obtained in experimental evaluations demonstrates the efficacy of the proposed approach. Bharath Vishwanath, Tejaswi Nanjundaswamy, Sina Zamani, Kenneth Rose |
ICASSP | 1 |
| 2017 | Rotational motion model for temporal prediction in 360 video codingabstractThe recent boom in the field of virtual and augmented reality has dramatically increased the prevalence of spherical video. Given the enormous amount of data consumed by spherical video, it is critical to achieve efficient compression for storage and transmission. Prevalent approaches simply project (via different geometries) the spherical video onto planes for processing with traditional 2D video coding standards. However, such approaches are significantly sub-optimal as standard video coders only allow for block translations in the critical tool of motion compensated prediction, which is incompatible with the expected motion in projected spherical video. Specifically, the effective sampling density varies over the sphere and the resulting locally varying warping yields complex non-linear motion in the projected domain. Hence, translation in the projected domain does not preserve an object's shape and size on the sphere, and its corresponding motion vector does not have a useful physical interpretation. Instead, we propose to characterize the motion directly on the sphere with a rotational motion model, specifically, in terms of sphere rotations along geodesics. This model preserves object shape and size on the sphere. A motion vector in this model implicitly specifies an axis of rotation and the degree of rotation about that axis, to convey the actual motion of objects on the sphere. Complementary to the novel motion model, we further propose an effective motion search technique that is tailored to the sphere's geometry. Experimental results demonstrate that the proposed framework achieves significant gains over prevalent motion models, across various projection geometries. Bharath Vishwanath, Tejaswi Nanjundaswamy, Kenneth Rose |
MMSP | 1 |
| 2017 | An evaluation framework for 360-degree video compressionabstract360-degree video is emerging as a new way of offering immersive visual experience. 360-degree video can be viewed on dedicated head mounted devices as well as on conventional 2D displays. Due to increased resolution to support wide field of view, efficient compression of 360-degree video becomes crucial. Whereas it is of significant interest to evaluate how different projection formats impact the compression efficiency of 360-degree video, a main technical challenge is that the input video is captured in a given native projection format, and that the native format has an obvious edge over other projection formats. In this paper, an evaluation framework is proposed to reduce the bias towards the native projection format when comparing different projection formats and their impact on 360-degree video compression. Additionally, a quality metric called area weighted spherical PSNR (AW-SPSNR) is proposed for objective 360-degree video quality evaluation. The proposed evaluation framework was included in the common test conditions defined for the exploration work of 360-degree video compression technologies under the Joint Video Exploration Team (JVET). Xiaoyu Xiu, Yuwen He, Yan Ye 0003, Bharath Vishwanath |
VCIP | 4 |