VLDB 2026 Research / reviewers in the wild / expert
Vignesh V. Menon
dblp:298/3337
· DBLP profile ↗
27ranked-venue papers
17as first author
27since 2021 · last 2026
0000-0003-1454-6146ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 17 first-author · 26 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Content-Driven Frame-Level Bit Prediction for Rate Control in Versatile Video Codingabstract3911 Amritha Premkumar, Prajit T. Rajendran, Vignesh V. Menon, Christian Herglotz |
ISCAS | 3 |
| 2025 | Machine Learning-Based Decoding Energy Modeling for VVC StreamingabstractEfficient video streaming requires jointly optimizing encoding parameters (bitrate, resolution, compression efficiency) and decoding constraints (computational load, energy consumption) to balance quality and power efficiency, particularly for resource-constrained devices. However, hardware heterogeneity, including differences in CPU/GPU architectures, thermal management, and dynamic power scaling, makes absolute energy models unreliable, particularly for predicting decoding consumption. This paper introduces the relative decoding energy index (RDEI), a metric that normalizes decoding energy consumption against a baseline encoding configuration, eliminating device-specific dependencies to enable cross-platform comparability and guide energyefficient streaming adaptations. We use a dataset of 1000 realistic video sequences to extract complexity features capturing spatial and temporal variations, employ Versatile Video Coding (VVC) open-source toolchain using VVenC/VVdeC with various resolutions, framerate, encoding preset and quantization parameter (QP) sets, and model RDEI using Random Forest (RF), XGBoost, Linear Regression (LR), and Shallow Neural Networks (NN) for decoding energy prediction. Experimental results demonstrate that RDEI-based predictions provide accurate decoding energy estimates across different hardware, ensuring cross-device comparability in VVC streaming. Reza Farahani, Vignesh V. Menon, Christian Timmerer |
ICIP | 2 |
| 2024 | Fast Constant-Quality Video Encoding Using VVENC With Rate Capping Based On Pre-Analysis StatisticsabstractVVenC, an open Versatile Video Coding (VVC) encoder, has recently been equipped with rate capping functionality in its two-pass rate control modes, providing constrained variable bitrate coding governed by target rate and maximum rate parameters. This paper reports on implementations and evaluation results of straightforward extensions to VVenC which enable the use of the maximum rate parameter also in the single-pass fixed-QP modes, controlled by a base quantization parameter (QP) instead of a target rate. The rate capping in the fixed-QP mode is achieved, with sufficient accuracy, by evaluating only already calculated pre-processing statistics, thereby avoiding increases in encoder runtime. This encoding mode, given that it supports visual quality optimizations such as XPSNR based block-wise perceptual QP adaptation, can be considered a rate capped constant-quality mode, which was missing in VVenC and which is an interesting configuration for video streaming. Christian R. Helmrich, Valeri George, Vignesh V. Menon, Adam Wieckowski, Benjamin Bross, Detlev Marpe |
ICIP | 3 |
| 2024 | Convex-Hull Estimation using Xpsnr for Versatile Video CodingabstractAs adaptive streaming becomes crucial for delivering high-quality video content across diverse network conditions, accurate metrics to assess perceptual quality are essential. This paper explores using the eXtended Peak Signal-to-Noise Ratio (XPSNR) metric as an alternative to the popular Video Multimethod Assessment Fusion (VMAF) metric for determining optimized bitrate-resolution pairs in the context of Versatile Video Coding (VVC). Our study is rooted in the observation that XPSNR shows a superior correlation with subjective quality scores for VVC-coded Ultra-High Definition (UHD) content compared to VMAF. We predict the average XPSNR of VVC-coded bitstreams using spatiotemporal complexity features of the video and the target encoding configuration and then determine the convex-hull online. On average, the proposed convex-hull using XPSNR (VEXUS) achieves an overall quality improvement of 5.84 dB PSNR and 0.62 dB XPSNR while maintaining the same bitrate, compared to the default UHD encoding using the VVenC encoder, accompanied by an encoding time reduction of 44.43% and a decoding time reduction of 65.46%. This shift towards XPSNR as a guiding metric shall enhance the effectiveness of adaptive streaming algorithms, ensuring an optimal balance between bitrate efficiency and perceptual fidelity with advanced video coding standards. Vignesh V. Menon, Christian R. Helmrich, Adam Wieckowski, Benjamin Bross, Detlev Marpe |
ICIP | 1 |
| 2024 | Quality-Aware Dynamic Resolution Adaptation Framework for Adaptive Video StreamingabstractTraditional per-title encoding schemes aim to optimize encoding resolutions to deliver the highest perceptual quality for each representation. XPSNR is observed to correlate better with the subjective quality of VVC-coded bitstreams. Towards this realization, we predict the average XPSNR of VVC-coded bitstreams using spatiotemporal complexity features of the video and the target encoding configuration using an XGBoost-based model. Based on the predicted XPSNR scores, we introduce a Quality-Aware Dynamic Resolution Adaptation (QADRA) framework for adaptive video streaming applications, where we determine the convex-hull online. Furthermore, keeping the encoding and decoding times within an acceptable threshold is mandatory for smooth and energy-efficient streaming. Hence, QADRA determines the encoding resolution and quantization parameter (QP) for each target bitrate by maximizing XPSNR while constraining the maximum encoding and/ or decoding time below a threshold. QADRA implements a JND-based representation elimination algorithm to remove perceptually redundant representations from the bitrate ladder. QADRA is an open-source Python-based framework published under the GNU GPLv3 license. Amritha Premkumar, Prajit T. Rajendran, Vignesh V. Menon, Adam Wieckowski, Benjamin Bross, Detlev Marpe |
MMSys | 3 |
| 2024 | Video Super-Resolution for Optimized Bitrate and Green Online StreamingabstractConventional per-title encoding schemes strive to optimize encoding resolutions to deliver the utmost perceptual quality for each bitrate ladder representation. Nevertheless, maintaining encoding time within an acceptable threshold is equally imperative in online streaming applications. Further-more, modern client devices are equipped with the capability for fast deep-learning-based video super-resolution (VSR) techniques, enhancing the perceptual quality of the decoded bitstream. This suggests that opting for lower resolutions in representations during the encoding process can curtail the overall energy consumption without substantially compromising perceptual quality. In this context, this paper introduces a video super-resolution-based latency-aware optimized bitrate encoding scheme (ViSOR) designed for online adaptive streaming applications. ViSOR determines the encoding resolution for each target bitrate, ensuring the highest achievable perceptual quality after VSR within the bound of a maximum acceptable latency. Random forest-based prediction models are trained to predict the perceptual quality after VSR and the encoding time for each resolution using the spatiotemporal features extracted for each video segment. Experimental results show that ViSOR targeting fast super-resolution convolutional neural network (FSRCNN) achieves an overall average bitrate reduction of 24.65% and 32.70% to maintain the same PSNR and VMAF, compared to the HTTP Live Streaming (HLS) bitrate ladder encoding of 4s segments using the x265 encoder, when the maximum acceptable latency for each representation is set as two seconds. Considering a just noticeable difference (JND) of six VMAF points, the average cumulative storage consumption and encoding energy for each segment is reduced by 79.32% and 68.21%, respectively, contributing towards greener streaming. Vignesh V. Menon, Prajit T. Rajendran, Amritha Premkumar, Benjamin Bross, Detlev Marpe |
PCS | 1 |
| 2024 | Towards ML-Driven Video Encoding Parameter Selection for Quality and Energy OptimizationabstractAs multimedia dominates Internet traffic, users seek a better Quality of Experience (QoE), often resulting in increased energy consumption and a higher carbon footprint. The increasing focus on sustainability underscores the critical need to balance energy consumption and QoE in video streaming. This paper proposes a modular architecture that refines video encoding parameters by assessing video complexity and encoding settings for the prediction of energy consumption and video quality (based on Video Multimethod Assessment Fusion (VMAF)) using lightweight XGBoost models trained on the multi-dimensional video compression dataset (MVCD). We apply Explainable AI (XAI) techniques to identify the critical encoding parameters that influence the energy consumption and video quality prediction models and then tune them using a weighting strategy between energy consumption and video quality. The experimental results confirm that applying a suitable weighting factor to energy consumption in the x265 encoder results in a 46 % decrease in energy consumption, with a 4-point drop in VMAF, staying below the Just Noticeable Difference (JND) threshold. Zoha Azimi Ourimi, Reza Farahani, Vignesh V. Menon, Christian Timmerer, Radu Prodan |
QoMEX | 3 |
| 2024 | Decoding Complexity-Rate-Quality Pareto-Front for Adaptive VVC StreamingabstractPareto-front optimization is crucial for addressing the multi-objective challenges in video streaming, enabling the identification of optimal trade-offs between conflicting goals such as bitrate, video quality, and decoding complexity. This paper explores the construction of efficient bitrate ladders for adaptive Versatile Video Coding (VVC) streaming, focusing on optimizing these trade-offs. We investigate various ladder construction methods based on Pareto-front optimization, including exhaustive Rate-Quality and fixed ladder approaches. We propose a joint decoding time-rate-quality Pareto-front, providing a comprehensive framework to balance bitrate, decoding time, and video quality in video streaming. This allows streaming services to tailor their encoding strategies to meet specific requirements, prioritizing low decoding latency, bandwidth efficiency, or a balanced approach, thus enhancing the overall user experience. The experimental results confirm and demonstrate these opportunities for navigating the decoding time-rate-quality space to support various use cases. For example, when prioritizing low decoding latency, the proposed method achieves a decoding time reduction of 14.86 % while providing Bjøntegaard delta rate savings of 4.65 % and 0.32 dB improvement in the eXtended Peak Signal-to-Noise Ratio (XPSNR)-Rate domain over the traditional fixed ladder solution. Vignesh V. Menon, Adam Wieckowski, Benjamin Bross, Detlev Marpe |
VCIP | 2 |
| 2024 | Energy-Quality-aware Variable Framerate Pareto-Front for Adaptive Video StreamingabstractOptimizing framerate for a given bitrate-spatial resolution pair in adaptive video streaming is essential to maintain perceptual quality while considering decoding complexity. Low framerates at low bitrates reduce compression artifacts and decrease decoding energy. We propose a novel method, Decoding-complexity aware Framerate Prediction (DECODRA), which employs a Variable Framerate Pareto-front approach to predict an optimized framerate that minimizes decoding energy under quality degradation constraints. DECODRA dynamically adjusts the framerate based on current bitrate and spatial resolution, balancing trade-offs between framerate, perceptual quality, and decoding complexity. Extensive experimentation with the Inter-4K dataset demonstrates DECODRA’s effectiveness, yielding an average decoding energy reduction of up to 13.45 %, with minimal VMAF reduction of 0.33 points at a low-quality degradation threshold, compared to the default 60 fps encoding. Even at an aggressive threshold, DECODRA achieves significant energy savings of 13.45 % while only reducing VMAF by 2.11 points. In this way, DECODRA extends mobile device battery life and reduces the energy footprint of streaming services by providing a more energy-efficient video streaming pipeline. Prajit T. Rajendran, Samira Afzal, Vignesh V. Menon, Christian Timmerer |
VCIP | 3 |
| 2024 | JND-Aware Two-Pass Per-Title Encoding Scheme for Adaptive Live StreamingabstractAdaptive live video streaming applications utilize a predefined collection of bitrate-resolution pairs, known as abitrate ladder, for simplicity and efficiency, eliminating the need for additional run-time to determine the optimal pairs during the live streaming session. These applications do not incorporate two-pass encoding methods due to increased latency. However, an optimized bitrate ladder could result in lower storage and delivery costs and improvedQuality of Experience(QoE). This paper presents a Just Noticeable Difference (JND)-aware constrained Variable Bitrate (cVBR) Two-pass Per-title encoding Scheme (JTPS) designed specifically for live video streaming. JTPS predicts a content- and JND-aware bitrate ladder using low-complexity features based onDiscrete Cosine Transform(DCT) energy and optimizes the constant rate factor (CRF) for each representation using random forest-based models. The effectiveness of JTPS is demonstrated using the open source video encoder x265, with an average bitrate reduction of 18.80% and 32.59% for the same PSNR and VMAF, respectively, compared to the standardHTTP Live Streaming(HLS) bitrate ladder using Constant Bitrate (CBR) encoding. The implementation of JTPS also resulted in a 68.96% reduction in storage space and an 18.58% reduction in encoding time for a JND of six VMAF points. Vignesh V. Menon, Prajit T. Rajendran, Christian Feldmann, Klaus Schöffmann, Mohammed Ghanbari 0001, Christian Timmerer |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | All-Intra Rate Control Using Low Complexity Video Features for Versatile Video CodingabstractVersatile Video Coding (VVC) allows for large compression efficiency gains over its predecessor, High Efficiency Video Coding (HEVC). The added efficiency comes at the cost of increased runtime complexity, especially for encoding. It is thus highly relevant to explore all available runtime reduction options. This paper proposes a novel first pass for two-pass rate control in all-intra configuration, using low-complexity video analysis and a Random Forest (RF)-based machine learning model to derive the data required for driving the second pass. The proposed method is validated using VVenC, an open and optimized VVC encoder. Compared to the default two-pass rate control algorithm in VVenC, the proposed method achieves around 32% reduction in encoding time for the preset faster, while on average only causing 2% BD-rate increase and achieving similar rate control accuracy. Vignesh V. Menon, Anastasia Henkel, Prajit T. Rajendran, Christian R. Helmrich, Adam Wieckowski, Benjamin Bross, Christian Timmerer, Detlev Marpe |
ICIP | 1 |
| 2023 | Optimizing Video Streaming for Sustainability and Quality: The Role of Preset Selection in Per-Title EncodingabstractHTTP Adaptive Streaming (HAS) methods divide a video into smaller segments, encoded at multiple pre-defined bitrates to construct a bitrate ladder. Bitrate ladders are usually optimized per title over several dimensions, such as bitrate, resolution, and framerate. This paper adds a new dimension to the bitrate ladder by considering the energy consumption of the encoding process. Video encoders often have multiple pre-defined presets to balance the trade-off between encoding time, energy consumption, and compression efficiency. Faster presets disable certain coding tools defined by the codec to reduce the encoding time at the cost of reduced compression efficiency. Firstly, this paper evaluates the energy consumption and compression efficiency of different x265 presets for 500 video sequences. Secondly, optimized presets are selected for various representations in a bitrate ladder based on the results to guarantee a minimal drop in video quality while saving energy. Finally, a new per title model, which optimizes the trade-off between compression efficiency and energy consumption, is proposed. The experimental results show that decreasing the VMAF score by 0.15 and 0.39 while choosing an optimized preset results in encoding energy savings of 70% and 83%, respectively. Hadi Amirpour, Vignesh V. Menon, Samira Afzal, Radu Prodan, Christian Timmerer |
ICME | 2 |
| 2023 | Just Noticeable Difference-Aware Per-Scene Bitrate-Laddering for Adaptive Video StreamingabstractIn video streaming applications, a fixed set of bitrate-resolution pairs (known as a bitrate ladder) is typically used during the entire streaming session. However, an optimized bitrate ladder per scene may result in (i) decreased storage or delivery costs or/and (ii) increased Quality of Experience. This paper introduces a Just Noticeable Difference (JND)-aware perscene bitrate ladder prediction scheme (JASLA) for adaptive video-on-demand streaming applications. JASLA predicts jointly optimized resolutions and corresponding constant rate factors (CRFs) using spatial and temporal complexity features for a given set of target bitrates for every scene, which yields an efficient constrained Variable Bitrate encoding. Moreover, bitrate-resolution pairs that yield distortion lower than one JND are eliminated. Experimental results show that, on average, JASLA yields bitrate savings of 34.42% and 42.67% to maintain the same PSNR and VMAF, respectively, compared to the reference HTTP Live Streaming (HLS) bitrate ladder Constant Bitrate encoding using x265 HEVC encoder, where the maximum resolution of streaming is Full HD (1080p). Moreover, a 54.34% average cumulative decrease in storage space is observed. Vignesh V. Menon, Prajit T. Rajendran, Hadi Amirpour, Patrick Le Callet, Christian Timmerer |
ICME | 1 |
| 2023 | Energy-Efficient Multi-Codec Bitrate-Ladder Estimation for Adaptive Video StreamingabstractWith the emergence of multiple modern video codecs, streaming service providers are forced to encode, store, and transmit bitrate ladders of multiple codecs separately, consequently suffering from additional energy costs for encoding, storage, and transmission. To tackle this issue, we introduce an online energy-efficient Multi-Codec Bitrate ladder Estimation scheme (MCBE) for adaptive video streaming applications. In MCBE, quality representations within the bitrate ladder of new-generation codecs (e.g., High Efficiency Video Coding (HEVC), Alliance for Open Media Video 1 (AV1)) that lie below the predicted rate-distortion curve of the Advanced Video Coding (AVC) codec are removed. Moreover, perceptual redundancy between representations of the bitrate ladders of the considered codecs is also minimized based on a Just Noticeable Difference (JND) threshold. Therefore, random forest-based models predict the VMAF score of bitrate ladder representations of each codec. In a live streaming session where all clients support the decoding of AVC, HEVC, and AV1, MCBE achieves impressive results, reducing cumulative encoding energy by 56.45%, storage energy usage by 94.99%, and transmission energy usage by 77.61% (considering a JND of six VMAF points). These energy reductions are in comparison to a baseline bitrate ladder encoding based on current industry practice. Vignesh V. Menon, Reza Farahani, Prajit T. Rajendran, Samira Afzal, Klaus Schöffmann, Christian Timmerer |
VCIP | 1 |
| 2023 | EMES: Efficient Multi-encoding Schemes for HEVC-based Adaptive Bitrate StreamingabstractIn HTTP Adaptive Streaming (HAS), videos are encoded at multiple bitrates and spatial resolutions ( i.e. , representations ) to adapt to the heterogeneity of network conditions, device attributes, and end-user preferences. Encoding the same video segment at multiple representations increases costs for content providers. State-of-the-art multi-encoding schemes improve the encoding process by utilizing encoder analysis information from already encoded representation(s) to reduce the encoding time of the remaining representations. These schemes typically use the highest bitrate representation as the reference to accelerate the encoding of the remaining representations. Nowadays, most streaming services utilize cloud-based encoding techniques, enabling a fully parallel encoding process to reduce the overall encoding time. The highest bitrate representation has the highest encoding time than the other representations. Thus, utilizing it as the reference encoding is unfavorable in a parallel encoding setup as the overall encoding time is bound by its encoding time. This paper provides a comprehensive study of various multi-rate and multi-encoding schemes in both serial and parallel encoding scenarios. Furthermore, it introduces novel heuristics to limit the Rate Distortion Optimization (RDO) process across various representations. Based on these heuristics, three multi-encoding schemes are proposed, which rely on encoder analysis sharing across different representations: (i) optimized for the highest compression efficiency , (ii) optimized for the best compression efficiency-encoding time savings trade-off , and (iii) optimized for the best encoding time savings . Experimental results demonstrate that the proposed multi-encoding schemes (i) , (ii) , and (iii) reduce the overall serial encoding time by 34.71%, 45.27%, and 68.76% with a 2.3%, 3.1%, and 4.5% bitrate increase to maintain the same VMAF, respectively compared to stand-alone encodings. The overall parallel encoding time is reduced by 22.03%, 20.72%, and 76.82% compared to stand-alone encodings for schemes (i) , (ii) , and (iii) , respectively. Vignesh V. Menon, Hadi Amirpour, Mohammed Ghanbari 0001, Christian Timmerer |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2022 | CODA: Content-aware Frame Dropping Algorithm for High Frame-rate Video StreamingabstractUltra High Definition Television (UHDTV) offers a better immersive audiovisual experience than HDTV by improving the aesthetic sense of the content [1]. How-ever, it may lead to an increase of both encoding time complexity and compression artifacts at lower bitrates. To address this challenge, a low-latency pre-processing algorithm named COntent-aware frame Dropping Algorithm (CODA) is proposed to predict the optimized framerate per video segment in streaming scenarios. The optimized framerate$(\hat{f})$for every video segment at each target bitrate is modelled as an exponential decay (increasing) function whose decay rate is directly proportional to the temporal characteristics$(h)$[2] [3] of the video and the target bitrate$(b)$, and inversely proportional to the spatial characteristics$(E)$of the video. The encoding is carried out with the predicted framerate, saving encoding time and improving visual quality at lower bitrates. At the decoder side, the video is upscaled in the temporal domain to the original framerate$(f_{max})$for display. Vignesh V. Menon, Hadi Amirpour, Mohammed Ghanbari 0001, Christian Timmerer |
DCC | 1 |
| 2022 | OPTE: Online Per-Title Encoding for Live Video StreamingabstractCurrent per-title encoding schemes encode the same video content at various bitrates and spatial resolutions to find an optimized bitrate ladder for each video content in Video on Demand (VoD) applications. However, in live streaming applications, a bitrate ladder with fixed bitrate-resolution pairs is used to avoid the additional latency caused to find optimum bitrate-resolution pairs for every video content. This paper introduces an online per-title encoding scheme (OPTE) for live video streaming applications. In this scheme, each target bitrate’s optimal resolution is predicted from any pre-defined set of resolutions using Discrete Cosine Transform (DCT)-energy-based low-complexity spatial and temporal features for each video segment. Experimental results show that, on average, OPTE yields bitrate savings of 20.45% and 28.45% to maintain the same PSNR and VMAF, respectively, compared to a fixed bitrate ladder scheme (as adopted in current live streaming deployments) without any noticeable additional latency in streaming. Vignesh V. Menon, Hadi Amirpour, Mohammed Ghanbari 0001, Christian Timmerer |
ICASSP | 1 |
| 2022 | ETPS: Efficient Two-Pass Encoding Scheme for Adaptive Live StreamingabstractIn two-pass encoding, also known as multi-pass encoding, the input video content is analyzed in the first-pass to help the second-pass encoding utilize better encoding decisions and improve overall compression efficiency. In live streaming applications, a single-pass encoding scheme is mainly used to avoid the additional first-pass encoding run-time to analyze the complexity of every video content. This paper introduces an Efficient low-latency Two-Pass encoding Scheme (ETPS) for live video streaming applications. In this scheme, Discrete Cosine Transform (DCT)-energy-based low-complexity spatial and temporal features for every video segment are extracted in the first-pass to predict each target bitrate’s optimal constant rate factor (CRF) for the second-pass constrained variable bitrate (cVBR) encoding. Experimental results show that, on average, ETPS compared to a traditional two-pass average bitrate encoding scheme yields encoding time savings of 43.78% without any noticeable drop in compression efficiency. Additionally, compared to a single-pass constant bitrate (CBR) encoding, it yields bitrate savings of 10.89% and 8.60% to maintain the same PSNR and VMAF, respectively. Vignesh V. Menon, Hadi Amirpour, Mohammed Ghanbari 0001, Christian Timmerer |
ICIP | 1 |
| 2022 | Perceptually-Aware Per-Title Encoding for Adaptive Video StreamingabstractIn live streaming applications, a fixed set of bitrate-resolution pairs (known as bitrate ladder) is used for simplicity and efficiency to avoid the additional encoding run-time required to find optimum resolution-bitrate pairs for every video content. However, an optimized bitrate ladder may result in (i) decreased storage or delivery costs or/and (ii) increased Quality of Experience (QoE). This paper introduces a perceptually-aware per-title encoding (PPTE) scheme for video streaming applications. In this scheme, optimized bitrate-resolution pairs are predicted online based on Just Noticeable Difference (JND) in quality perception to avoid adding perceptually similar representations in the bitrate ladder. To this end, Discrete Cosine Transform (DCT)-energy-based low-complexity spatial and temporal features for each video segment are used. Experimental results show that, on average, PPTE yields bitrate savings of 16.47% and 27.02% to maintain the same PSNR and VMAF, respectively, compared to the reference HTTP Live Streaming (HLS) bitrate ladder without any noticeable additional latency in streaming accompanied by a 30.69% cumulative decrease in storage space for various representations. Vignesh V. Menon, Hadi Amirpour, Mohammed Ghanbari 0001, Christian Timmerer |
ICME | 1 |
| 2022 | Video Coding Enhancements for HTTP Adaptive StreamingabstractRapid growth in multimedia streaming traffic over the Internet motivates the research and further investigation of the video coding performance of such services in terms of speed and Quality of Experience (QoE). HTTP Adaptive Streaming (HAS) is today's de-facto standard to deliver clients the highest possible video quality. In HAS, the same video content is encoded at multiple bitrates, resolutions, framerates, and coding formats called representations. This study aims to (i) provide fast and compression-efficient multi-bitrate, multi-resolution representations, (ii) provide fast and compression-efficient multi-codec representations, (iii) improve the encoding efficiency of Video on Demand (VoD) streaming using content-adaptive encoding optimizations, and (iv) provide encoding schemes with optimizations per-title for live streaming applications to decrease the storage or delivery costs or/and increase QoE. Vignesh V. Menon |
ACM Multimedia | 1 |
| 2022 | Light-weight Video Encoding Complexity Prediction using Spatio Temporal FeaturesabstractThe increasing demand for high-quality and low-cost video streaming services calls for the prediction of video encoding complexity. The prior prediction of video encoding complexity including encoding time and bitrate predictions are used to allocate resources and set optimized parameters for video encoding effectively. In this paper, a light-weight video encoding complexity prediction (VECP) scheme that predicts the encoding bitrate and the encoding time of video with high accuracy is proposed. Firstly, low-complexity Discrete Cosine Transform (DCT)-energy-based features, namely spatial complexity, temporal complexity, and brightness of videos are extracted, which can efficiently represent the encoding complexity of videos. The latent vectors are also extracted from a Convolutional Neural Network (CNN) with MobileNet as the backend to obtain additional features from representative frames of each video to assist the prediction process. The extreme gradient boosting (XGBoost) regression algorithm is deployed to predict video encoding complexity using the extracted features. The experimental results demonstrate that VECP predicts the encoding bitrate with an error percentage of up to 3.47% and encoding time with an error percentage of up to 2.89%, but with a significantly low overall latency of 3.5 milliseconds per frame which makes it suitable for both Video on Demand (VoD) and live streaming applications. Hadi Amirpour, Prajit T. Rajendran, Vignesh V. Menon, Mohammed Ghanbari 0001, Christian Timmerer |
MMSP | 3 |
| 2022 | VCD: video complexity datasetabstractThis paper provides an overview of the open Video Complexity Dataset (VCD) which comprises 500 Ultra High Definition (UHD) resolution test video sequences. These sequences are provided at 24 frames per second (fps) and stored online in losslessly encoded 8-bit 4:2:0 format. In this paper, all sequences are characterized by spatial and temporal complexities, rate-distortion complexity, and encoding complexity with the x264 AVC/H.264 and x265 HEVC/H.265 video encoders. The dataset is tailor-made for cutting-edge multimedia applications such as video streaming, two-pass encoding, per-title encoding, scene-cut detection, etc. Evaluations show that the dataset includes diversity in video complexities. Hence, using this dataset is recommended for training and testing video coding applications. All data have been made publicly available as part of the dataset, which can be used for various applications. Hadi Amirpour, Vignesh V. Menon, Samira Afzal, Mohammed Ghanbari 0001, Christian Timmerer |
MMSys | 2 |
| 2022 | VCA: video complexity analyzerabstractFor online analysis of the video content complexity in live streaming applications, selecting low-complexity features is critical to ensure low-latency video streaming without disruptions. To this light, for each video (segment), two features, i.e., the average texture energy and the average gradient of the texture energy, are determined. A DCT-based energy function is introduced to determine the block-wise texture of each frame. The spatial and temporal features of the video (segment) are derived from this DCT-based energy function. The Video Complexity Analyzer (VCA) project aims to provide an efficient spatial and temporal complexity analysis of each video (segment) which can be used in various applications to find the optimal encoding decisions. VCA leverages some of the x86 Single Instruction Multiple Data (SIMD) optimizations for Intel CPUs and multi-threading optimizations to achieve increased performance. VCA is an open-source library published under the GNU GPLv3 license. Vignesh V. Menon, Christian Feldmann, Hadi Amirpour, Mohammed Ghanbari 0001, Christian Timmerer |
MMSys | 1 |
| 2022 | Content-adaptive Encoder Preset Prediction for Adaptive Live StreamingabstractIn live streaming applications, a fixed set of bitrate-resolution pairs (known as bitrate ladder) is generally used to avoid additional pre-processing run-time to analyze the complexity of every video content and determine the optimized bitrate ladder. Furthermore, live encoders use the fastest available preset for encoding to ensure the minimum possible latency in streaming. For live encoders, it is expected that the encoding speed is equal to the video framerate. An optimized encoding preset may result in (i) increased Quality of Experience (QoE) and (ii) improved CPU utilization while encoding. In this light, this paper introduces a Content-Adaptive encoder Preset prediction Scheme (CAPS) for adaptive live video streaming applications. In this scheme, the encoder preset is determined using Discrete Cosine Transform (DCT)-energy-based low-complexity spatial and temporal features for every video segment, the number of CPU threads allocated for each encoding instance, and the target encoding speed. Experimental results show that ChPS yields an overall quality improvement of 0.83 dB PSNR and 3.81 VMAF with the same bitrate, compared to the fastest preset encoding of the HTTP Live Streaming (HLS) bitrate ladder using $\times265$ HEVC open-source encoder. This is achieved by maintaining the desired encoding speed and reducing CPU idle time. Vignesh V. Menon, Hadi Amirpour, Prajit T. Rajendran, Mohammed Ghanbari 0001, Christian Timmerer |
PCS | 1 |
| 2021 | Efficient Content-Adaptive Feature-Based Shot Detection for HTTP Adaptive StreamingabstractVideo delivery over the Internet has been becoming a commodity in recent years, owing to the widespread use of Dynamic Adaptive Streaming over HTTP (DASH). The DASH specification defines a hierarchical data model for Media Presentation Descriptions (MPDs) in terms of segments. This paper focuses on segmenting video into multiple shots for encoding in Video on Demand (VoD) HTTP Adaptive Streaming (HAS) applications. Therefore, we propose a novel Discrete Cosine Transform (DCT) feature-based shot detection and successive elimination algorithm for shot detection and compare it against the default shot detection algorithm of the x265 implementation of the High Efficiency Video Coding (HEVC) standard. Our experimental results demonstrate that our proposed feature-based pre-processor has a recall rate of 25% and an F-measure of 20% greater than the benchmark algorithm for shot detection. Vignesh V. Menon, Hadi Amirpour, Mohammed Ghanbari 0001, Christian Timmerer |
ICIP | 1 |
| 2021 | INCEPT: Intra CU Depth Prediction for HEVCabstractHigh Efficiency Video Coding (HEVC) improves the encoding efficiency by utilizing sophisticated tools such as flexible Coding Tree Units (CTUs) partitioning. The Coding Units (CUs) can be split recursively into four equally sized CUs ranging from 64×64 to 8×8 pixels. At each depth level (or CU size), intra prediction via exhaustive mode search was exploited in HEVC to improve the encoding efficiency and result in a very high encoding time complexity. This paper proposes an Intra CU Depth Prediction (INCEPT) algorithm, which limits Rate-Distortion Optimization (RDO) for each CTU in HEVC by utilizing the spatial correlation with the neighboring CTUs, which is computed using a DCT energy-based feature. Thus, INCEPT reduces the number of candidate CU sizes required to be considered for each CTU in HEVC intra coding. Experimental results show that the INCEPT algorithm achieves a better trade-off between the encoding efficiency and encoding time saving (i.e., BDR/∆T) than the benchmark algorithms. While BDR/∆T is 12.35% and 9.03% for the benchmark algorithms, it is 5.49% for the proposed algorithm. As a result, INCEPT achieves a 23.34% reduction in encoding time on average while incurring only a 1.67% increase in bitrate than the original coding in the x265 HEVC open-source encoder. Vignesh V. Menon, Hadi Amirpour, Christian Timmerer, Mohammed Ghanbari 0001 |
MMSP | 1 |
| 2021 | Efficient Multi-Encoding Algorithms for HTTP Adaptive Bitrate StreamingabstractSince video accounts for the majority of today's internet traffic, the popularity of HTTP Adaptive Streaming (HAS) is increasing steadily. In HAS, each video is encoded at multiple bitrates and spatial resolutions (i.e., representations) to adapt to a heterogeneity of network conditions, device characteristics, and end-user preferences. Most of the streaming services utilize cloud-based encoding techniques which enable a fully parallel encoding process to speed up the encoding and consequently to reduce the overall time complexity. State-of-the-art approaches further improve the encoding process by utilizing encoder analysis information from already encoded representation(s) to improve the encoding time complexity of the remaining representations. In this paper, we investigate various multi-encoding algorithms (i.e., multi-rate and multi-resolution) and propose novel multi-encoding algorithms for large-scale HTTP Adaptive Streaming deployments. Experimental results demonstrate that the proposed multi-encoding algorithm optimized for the highest compression efficiency reduces the overall encoding time by 39% with a 1.5% bitrate increase compared to stand-alone encodings. Its optimized version for the highest time savings reduces the overall encoding time by 50% with a 2.6% bitrate increase compared to standalone encodings. Vignesh V. Menon, Hadi Amirpour, Christian Timmerer, Mohammed Ghanbari 0001 |
PCS | 1 |