Jarno Vanne

dblp:03/16 · DBLP profile ↗
← Back
81ranked-venue papers
7as first author
34since 2021 · last 2026
0000-0002-7944-1938ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 60 · 6 first-author · 28 since 2021Systems, architecture and hardware · 21 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2026 UVG-VCM: Benchmarking Dataset for Machine-Oriented Visual Data Compression
Tero Partanen, Miro Anttila, Rudolf Kortelahti, Guillaume Gautier, Alexandre Mercat, Jarno Vanne
QoMEX6
2026 UVG-GS: Human-Centric Gaussian Splats Dataset for Visual Volumetric Compression
Uyen Phan, Alexandre Mercat, Patrice Rondao-Alface, Louis Fréneau, Jarno Vanne, Guillaume Gautier
QoMEX5
2025 ShapeFuture - Technical Progress After Year 1
abstract
ShapeFuture will drive innovation in fundamental Electronic Components and Systems (ECS) that are essential for robust, powerful, fail-operational and integrated perception, cognition, AI-enabled decision making, resilient automation and computing, as well as communications, for highly automated vehicles. The overarching vision of ShapeFuture is to bring ECS Innovation to the heart of Europe’s Mobility Transformation, thereby elevating sovereignty by perfecting programmable ECS solutions for intelligent, safe, connected, and highly automated vehicles. In this paper, we detail not only the vision and mission of the ShapeFuture project, but we also showcase the results achieved during the first year.
Norbert Druml, Martin Gschwandtner, Mayeul Jeannin, Rainer Matischek, Edgars Lielamurs, Maksis Celitans, Kaspars Ozols, Nurullah Demiralay, Besir Tayfur, Ismail Sinan Gulbas, Nadir Kucuk, Isa Kiyat, Yahya Nasolo, Jens U. Brandt, Noah Christoph Pütz, Thomas Bartz-Beielstein, Jose Isola, Nikola Mandic, Francesca Flamigni, Alexander Kuehhas, Gianluca Brilli, Paolo Burgio, Giacomo Paolieri, Jorge Villagra, Álvaro Flores Cueto, José Antonio Sánchez, Jacopo Sini, Massimo Violante, Lorenzo Giraudi, Paolo Santero, Uwe Kölbel, Moritz Schaffenroth, Panu Sjövall, Jarno Vanne, Morten Larsen, Nergis Gizem Yilmaz, Ziya Uygar Yengin, George Dimitrakopoulos 0001
DSD34
2025 Live Demonstration: VVC Intra-Coding Visualizer
abstract
This paper presents a live demonstrator that is designed to visualize the encoding process of uvg266 VVC open-source intra encoder. The proposed visualizer provides real-time insights into the decision-making process of uvg266 by illustrating chosen block structures, modes, transforms, and parallelization strategies. The tool also facilitates interactive analysis through a user-friendly interface that allows users to dynamically adjust encoding tools and parameters on the fly. Researchers, developers, and students can use the tool to enhance their understanding of VVC coding flow and to support their optimization and educational efforts.
Joose Sainio, Alexandre Mercat, Jarno Vanne
ISCAS3
2025 UVG-CWI-DQPC: Dual-Quality Point Cloud Dataset for Volumetric Video Applications
abstract
Volumetric video is a key enabler of immersive extended reality (XR) experiences and is often represented using point clouds for their structural simplicity. However, capturing volumetric content through multi-view acquisition and depth sensing poses many challenges, such as occlusions and depth mismatches. To foster research in this field, we introduce a unique dual-quality point cloud dataset, named UVG-CWI-DQPC, which is designed to support the development of point cloud enhancement, compression, and quality assessment. Our dataset includes 12 dynamic sequences captured simultaneously by: 1) a high-end capture system producing high-fidelity point clouds with extensive processing; and 2) a consumer-grade capture system relying on affordable RGB-D cameras, lightweight processing, and open-source tools. For each sequence, our dataset provides ground-truth point clouds from the high-end capture system and raw RGB-D footage from the consumer-grade capture system, along with calibration data and tools for point cloud generation. This dual-quality setup enables direct comparison and benchmarking of algorithms for densification, occlusion removal, registration, and quality enhancement. Our dataset is publicly available under a permissive license to support reproducible research and standardization work in Moving Picture Experts Group (MPEG) and 3rd Generation Partnership Project (3GPP).
Guillaume Gautier, Xuemei Zhou, Jack Jansen 0001, Louis Fréneau, Marko Viitanen, Uyen Phan, Jani Käpylä, Irene Viola 0001, Alexandre Mercat, Pablo César, Jarno Vanne
ACM Multimedia12
2025 uvgVPCCenc: Practical Open-Source Encoder for Fast V-PCC Compression
abstract
Video-based Point Cloud Compression (V-PCC) standard offers state-of-the-art tools and efficiency for volumetric video compression. The V-PCC reference software, TMC2, is able to demonstrate the compression efficiency of V-PCC, but with computational complexity that is impractical for real-life applications. This paper introduces a novel open-source V-PCC encoder, named uvgVPCCenc, that brings V-PCC compression speed to a practical level. To reduce computational bottlenecks, uvgVPCCenc is implemented from the ground up in C++ with advanced multi-threading and streamlined algorithms. It also supports a flexible parameterization and different presets that facilitate its adaptation to various content and target bitrates. Our comparative evaluations show that uvgVPCCenc achieves substantial speedups over TMC2 while maintaining competitive rate-distortion-complexity tradeoff. To the best of our knowledge, uvgVPCCenc is the first open-source encoder explicitly designed for practical V-PCC compression, paving the way from theoretical research to practical volumetric video applications.
Louis Fréneau, Guillaume Gautier, Alexandre Mercat, Jarno Vanne
MMSys4
2025 Task Enhancement Tiles for Ultra Lightweight Post-processing in Visual Coding for Machines
abstract
The proliferation of automated visual analysis calls for compression methods tailored to the unique requirements of Video Coding for Machines (VCM). In this paper, we propose a computationally lightweight post-processing method that is based on a learned component referred to as a task enhancement tile (TET). A TET is spatially tiled over the reconstructed visual data and added to it element-wise. It only requires one addition per pixel in each color channel before the machine task can be applied. Our results with the VVC test model (VTM) demonstrate coding gains of up to 39.0% for object detection and 29.2% for instance segmentation on image datasets, while evaluation on a video dataset shows gains of up to 35.2% for object detection, relative to the VTM anchor. The proposed solution also offers extremely low computational cost, preservation of human-viewable content, full compliance with video coding standards, no requirement for side information transmission from encoder to decoder, and generalization across tasks, models, and encoding parameters.
Tero Partanen, Alban Marie, Rudolf Kortelahti, Alexandre Mercat, Jarno Vanne, Miska M. Hannuksela, Honglei Zhang 0001, Alireza Aminlou, Francesco Cricri
PCS5
2025 Lightweight Motion-Vector-Based Segmentation Mask Tracking for Saliency-Based Rate Control
abstract
In video encoding, saliency-based rate control improves coding efficiency without compromising perceptual quality by allocating more bits to visually important regions. While object detection can guide bit allocation through rectangular bounding boxes, segmentation masks offer a more accurate delineation of object boundaries. This paper seeks to reduce the significant computational burden of frame-by-frame instance segmentation by proposing a lightweight segmentation mask tracking scheme, in which motion vectors (MVs) from the video encoder are used to predict per-vertex displacements. Altogether, we propose two neural network designs for segmentation mask tracking: (1) a base tracker optimized for tracking accuracy; and (2) a lite tracker that balances accuracy and computational complexity. Our experimental results show that the base tracker attains 70-88% of the accuracy of frame-by-frame instance segmentation but achieves a 48× speedup and reduces computational complexity to 0.03% on CPU. For the lite tracker, the corresponding figures are 49×, 0.01%, and 67-88%. Despite tradeoffs in tracking accuracy, reducing complexity to a fraction makes our solution a viable option for practical applications.
Tero Partanen, Jarno Vanne, Alexandre Mercat
VCIP3
2025 High-Level Synthesis FPGA Implementation of Fractional Motion Estimation for HEVC and VVC
abstract
The rapid deployment of modern video coding standards underscores the need for hardware implementations that enable real-time coding, support interoperability across codecs, and are openly accessible to community. This paper presents the first known high-level synthesis (HLS) implementation of an accurate full-search fractional motion estimation (FME). The proposed FME core is released as open-source and is compatible with High Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC) standards. It implements 1) an accurate multiplierless constant multiplication (MCM) unit that performs quarter-pixel interpolation over 9 × 9 pixels at a time; 2) a transform-exempted sum of absolute transformed differences (TE-SATD) unit that is optimized for area and speed; and 3) fixed-point Lagrangian optimizations for rate-distortion optimization (RDO). On an Intel Arria 10 FPGA, the FME core consumes 122 kALUTs and operates at up to 210 MHz. Our profiling results show that a single FME core can support practical HEVC and VVC encoding of 2160p video at 30–120 fps, depending on the video content and encoder preset.
Jesse Smedberg, Panu Sjövall, Guillaume Gautier, Alexandre Mercat, Jarno Vanne
VCIP5
2024 Feasibility Study of Multi-Layer VVC Coding Scheme for Hybrid Machine-Human Consumption
abstract
The proliferation of machine vision applications necessitates developing more efficient visual data compression schemes for machine consumption. However, numerous automated use cases still require keeping humans in the loop, leading to the need for a machine-optimized video streaming with the option for human supervision. This paper investigates the feasibility of using the multi-layer coding approach of the emerging Versatile Video Coding (VVC) standard to create favorable conditions for hybrid machine-human consumption. We introduce a multi-layer coding scheme, where the base layer (BL) is optimized for machines and the enhancement layer (EL) complements the stream for human vision. Our results demonstrate that the bitrate of the proposed multi-layer stream (BL + EL) is, on average, 11% higher than that of a single-layer VVC. However, the more compact BL yields overall bandwidth savings as long as the EL is required less than 80% of the time.
Jaakko Laitinen, Tero Partanen, Alexandre Mercat, Jarno Vanne, Miska M. Hannuksela, Honglei Zhang 0001, Alireza Aminlou, Francesco Cricri
ICME4
2024 uvgComm: Open Software for Low-Latency Multi-party Video Communication
abstract
Effective and smooth video communication has become an integral tool for interaction in our modern society. The well-known WebRTC project is optimized for server-based communication architectures with less stringent latency requirements. This paper introduces an open-source, peer-to-peer (P2P) mesh-based application called uvgComm 1.0 for low-latency multi-party video communication. Our experimental results show that uvgComm attains 25% lower latency than server-based solutions at 219 ms. In addition, it offers advanced privacy protection mechanisms with end-to-end encryption and P2P mesh topology that delivers media without any media servers. The ambitious goal of uvgComm is to take P2P mesh-based conferencing architecture to the next level and make this architecture a serious competitor to server-based communication architectures.
Joni Räsänen, Heikki Tampio, Alexandre Mercat, Jarno Vanne
ACM Multimedia4
2024 uvg266: Open-Source VVC Intra Encoder
abstract
Versatile Video Coding (VVC/H.266) standard is the emerging successor to the widespread High Efficiency Video Coding (HEVC/H.265). This work introduces the latest version of our academic open-source VVC intra encoder called uvg266. It has been developed from our well-known Kvazaar HEVC encoder by introducing new VVC coding tools into carefully optimized and parallelized coding flow of Kvazaar. This paper outlines the design methodology and implementation aspects of all intra (AI) configuration of uvg266. The experimental results show that single-threaded uvg266 is more than twice as fast as the state-of-the-art VVenC encoder in all our test cases. In speed-optimized coding, the coding overhead of uvg266 is 21.7% but the gap narrows down to 2.6% in the rate distortion optimized coding case. uvg266 has almost linear speedup with core count up to 32 threads. The better scalability of uvg266 quadruples the speed over VVenC. Furthermore, single-threaded uvg266 is up to 380× as fast as VVC reference software VTM and the gap raises to over 11,000× with 32 threads. To the best of our knowledge, uvg266 is currently the fastest available open-source VVC intra software encoder.
Marko Viitanen, Joose Sainio, Kari Siivonen, Alexandre Mercat, Jarno Vanne
ACM Multimedia5
2024 Luma Range Scaling for Enhanced VVC Efficiency in Video Coding for Machines
abstract
Recent years have shown significant growth in video data traffic for machine vision applications, catalyzing new standardization efforts in video coding for machines (VCM). These activities focus on compressing images and videos for machine vision tasks, rather than for human viewing. In this work, we propose a novel method that scales down the luma range to enhance the coding efficiency of Versatile Video Coding (VVC) for machine consumption. This method results in a lower bitrate after encoding and has only minimal adverse effects on the accuracy of machine vision tasks. In our experiments, we down-scale the luma channel of the input video using luma-scaling factors from 0.2 to 0.9 and evaluate coding results with optional back-scaling to the original range before machine vision tasks. Our results with the VVC Test Model (VTM) demonstrate that the proposed technique achieves coding gain of up to 37.9%and 46.1% for the same object detection and tracking accuracy, respectively.
Tero Partanen, Alban Marie, Alexandre Mercat, Jarno Vanne, Miska M. Hannuksela, Honglei Zhang 0001, Alireza Aminlou, Francesco Cricri
MMSP4
2024 uvgRTP 3.0: Towards V3C Volumetric Video Communication
abstract
Low-latency volumetric video transport is a key enabling technology for more immersive communication applications. This paper presents the latest release of our open-source Real-time Transport Protocol (RTP) library called uvgRTP 3.0 that has been upgraded to support Visual Volumetric Video-based Coding (V3C) transmission in Video-based Point Cloud Compression (V-PCC) and MPEG Immersive Video (MIV) formats. uvgRTP 3.0 introduces 1) the V3C atlas RTP payload format; 2) two multiplexing methods to reduce port reservations during V3C transmission; and 3) improved packet reception with multithreading. Our performance results show that uvgRTP 3.0 can encrypt and transmit V3C bitstreams at 106 Mbit/s, with CPU core utilization of 26%, and achieving a round-trip latency of 2 ms in a local area network. The support for high-speed, encrypted V3C communication with the permissive BSD-license make uvgRTP 3.0 a potential transmission library for any industrial or academic volumetric communication system.
Heikki Tampio, Joni Räsänen, Marko Viitanen, Alexandre Mercat, Guillaume Gautier, Jarno Vanne
MMSys6
2024 Tool Space Exploration of the V-PCC Patch Generation for Practical Point Cloud Encoding
abstract
Video-based Point Cloud Compression (V-PCC) is the latest MPEG standard for visual volumetric data coding. V-PCC leverages the coding efficiency of well-established 2D video codecs by creating patches that link sub-regions of a point cloud to their 2D projections. This paper provides an in-depth analysis of the V-PCC patch generation process, which is the most compute-intensive part of the V-PCC standard. To the best of our knowledge, this is the first work to explore the patch generation tools individually and evaluate their impact on coding efficiency and complexity with different coding parameters. The main purpose of this characterization is to offer key insights into the V-PCC encoder design process and thereby foster the development of practical V-PCC encoders.
Louis Fréneau, Alexandre Mercat, Guillaume Gautier, Joose Sainio, Jarno Vanne
PCS5
2024 Real-Time Video-based Point Cloud Compression
abstract
This paper presents a demonstration setup for our open-source intra encoder called uvgVPCCenc, which is optimized for real-time Video-based Point Cloud Compression (V-PCC). uvgVPCCenc achieves an average encoding speed of 26 frames per second (fps) on an Intel i7-12700 CPU when encoding volumetric video sequences with up to 185 000 points per frame. It is shown to be 700 times as fast as TMC2 reference implementation for V-PCC. Our work is the first to demonstrate real-time intra V-PCC encoding on a consumer-grade desktop computer. It indicates that even the immense computational complexity of intra V-PCC encoding can be tackled for practical applications with effective design and optimization techniques.
Louis Fréneau, Guillaume Gautier, Heikki Tampio, Alexandre Mercat, Jarno Vanne
VCIP5
2024 Fast Machine Learning Aided Intra Mode Decision for Real-Time VVC Intra Coding
abstract
Reducing the huge computational complexity of intra mode decision is the key to real-time Video Coding (VVC). This paper proposes a fast intra mode decision scheme that takes advantage of lightweight machine learning (ML) models to classify intra modes into fifteen clusters. The cluster is further refined using one of the three proposed strategies to select the most optimal mode. Our experimental results with the fastest configuration of the practical uvg266 encoder show that the proposed methods yield a competitive rate-distortion-complexity trade-off over a conventional rough mode decision (RMD). To the best of our knowledge, this is the first work to successfully reduce the complexity of RMD in a practical VVC encoder with the use of ML techniques.
Joose Sainio, Baran Ataman, Alban Marie, Alexandre Mercat, Jarno Vanne
VCIP5
2024 Vectorized Angular Intra Prediction for Practical VVC Encoding
abstract
Versatile Video Coding (VVC) provides new coding tools for more efficient intra prediction but with a substantial increase in computational complexity. This paper introduces vectorized kernels for 8-bit angular intra prediction and position dependent intra prediction combination (PDPC), which are carefully optimized for all block sizes and prediction modes of VVC. The proposed kernels streamline the filtering process and utilize optimized memory access patterns. Our standalone tests show that the proposed vectorization achieves speedups of 6.68× for luma and 4.40× for chroma predictions over scalar implementations. Integrating these kernels into the practical uvg266 VVC encoder provides speedups of 1.07× in the slowest configuration and 1.68× in the fastest configuration. The reported speedups are obtained without any coding overhead, so the proposed vectorization plays an integral role in pursuing real-time VVC coding with high coding efficiency.
Kari Siivonen, Joose Sainio, Guillaume Gautier, Alexandre Mercat, Jarno Vanne
VCIP5
2023 RDO Candidate Selection for Maximizing Coding Efficiency in a Practical HEVC Encoder
abstract
High Efficiency Video Coding (HEVC) creates the conditions for economic video transmission and storage but making most of its compression potential calls for effective rate-distortion optimization (RDO) techniques in practical HEVC encoders. This paper explores the effectiveness of the following universally applicable RDO techniques: 1) rough mode decision for intra RDO candidate selection; 2) number of intra and inter RDO search candidates; and 3) accurate bit cost estimation in entropy coding. All these techniques are implemented into Kvazaar open-source HEVC encoder and altogether they improve the coding efficiency of Kvazaar veryslow preset by 5.7%, 4.0%, and 7.1% with PSNR, SSIM, and VMAF quality metrics, respectively. Even though the proposed techniques reduce the coding speed of Kvazaar to 0.64×, Kvazaar is still, on average, 2.16× as fast as the x265 encoder and attains 12.4%, 22.3%, and 4.7% better coding gain for the same PSNR, SSIM, and VMAF quality, respectively. These results let us conclude that Kvazaar is currently the leading practical open-source solution for high-quality HEVC encoding.
Joose Sainio, Alexandre Mercat, Jarno Vanne
ICASSP3
2023 FPGA-Accelerated HEVC Encoder for Energy-Efficient Multi-Access Edge Computing
abstract
High Efficiency Video Coding (HEVC) and Multi-access Edge Computing (MEC) technologies can make real-time streaming media services available to users with reasonable bandwidth, but the computational complexity of HEVC tends to lead to increased energy consumption in these schemes. In this paper, we investigate the energy saving opportunities of utilizing a field-programmable gate array (FPGA) based HEVC encoder in edge media servers and devices. In practice, we analyze the energy impact of migrating our Kvazaar software HEVC intra encoder to Intel Arria 10 PCIe FPGA(s) on two platforms: 1) Nokia Airframe Cloud Server with 2.4 GHz dual 14-core Intel Xeon processors and 2) an embedded Jetson AGX Orin board with 2.2 GHz 12-core ARM processor. According to our experiments, FPGA encoding on these two platforms saved 76% and 86% of the energy taken up by software only encoding on Airframe, respectively. These results indicate the potential of FPGA-based video encoder acceleration in future green MEC architectures.
Panu Sjövall, Alexandre Mercat, Jarno Vanne
ICIP3
2023 AVX2-Optimized Interpolation Filters for HEVC Inter Encoding
abstract
High Efficiency Video Coding (HEVC) sets the stage for economic video transmission and storage, but its inherent computational complexity calls for powerful implementations. This paper addresses the principal performance bottleneck of HEVC codecs by introducing AVX2-vectorized algorithms for HEVC interpolation filters. The proposed speed-up techniques include 1) a data permutation scheme for the horizontal interpolation stage; 2) a sliding window strategy for the vertical interpolation stage; 3) optimal usage of horizontal and vertical interpolation during fractional motion estimation; and 4) a lane-based approach to double the vector lengths from 128-bit legacy vector extensions to 256bits of AVX2. Our AVX2-optimized interpolation filters were benchmarked as part of the practical Kvazaar open-source HEVC encoder. On an Intel 8-core Xeon processor, they were shown to be 9.7 and 8.5 times as fast as scalar interpolation with the Kvazaar ultrafast and veryslow presets, respectively. In both cases, changing over from scalar to vectorized interpolation increases the coding speed of Kvazaar by more than 50%, which stresses the importance of interpolation optimizations in modern video encoders.
Alexandre Mercat, Ari Lemmetti, Joose Sainio, Jarno Vanne
ISCAS4
2023 Open-Source Toolkit for Live End-to-End 4K VVC Intra Coding
abstract
Versatile Video Coding (VVC/H.266) takes video coding to the next level by doubling the coding efficiency over its predecessors for the same subjective quality, but at the cost of immense coding complexity. Therefore, VVC calls for aggressively optimized codecs to make it feasible for live streaming media applications. This paper introduces the first public end-to-end (E2E) pipeline for live 4K30p VVC intra coding and streaming. The pipeline is made up of three open-source components: 1) uvg266 for VVC encoding; 2) uvgRTP for VVC streaming; and 3) OpenVVC for VVC decoding. The proposed setup is demonstrated with a proof-of-concept prototype that implements the encoder end on AMD ThreadRipper 2990WX and the decoder end on Nvidia Jetson AGX Orin. Our prototype is almost 34 000 times as fast as the corresponding E2E pipeline built around the VTM codec. Respectively, it achieves 3.3 times speedup without any significant coding overhead over the pipeline that utilizes the fastest possible configuration of the well-known VVenC/VVdeC codec. These results indicate that our prototype is currently the only viable open-source solution for live 4K VVC intra coding and streaming.
Marko Viitanen, Joose Sainio, Alexandre Mercat, Guillaume Gautier, Jarno Vanne, Ibrahim Farhat, Pierre-Loup Cabarat, Wassim Hamidouche, Daniel Ménard
MMSys5
2023 UVG-VPC: Voxelized Point Cloud Dataset for Visual Volumetric Video-based Coding
abstract
Point cloud compression has become a crucial factor in immersive visual media processing and streaming. This paper presents a new open dataset called UVG-VPC for the development, evaluation, and validation of MPEG Visual Volumetric Video-based Coding (V3C) technology. The dataset is distributed under its own non-commercial license. It consists of 12 point cloud test video sequences of diverse characteristics with respect to the motion, RGB texture, 3D geometry, and surface occlusion of the points. Each sequence is 10 seconds long and comprises 250 frames captured at 25 frames per second. The sequences are voxelized with a geometry precision of 9 to 12 bits, and the voxel color attributes are represented as 8-bit RGB values. The dataset also includes associated normals that make it more suitable for evaluating point cloud compression solutions. The main objective of releasing the UVG-VPC dataset is to foster the development of V3C technologies and thereby shape the future in this field.
Guillaume Gautier, Alexandre Mercat, Louis Fréneau, Mikko Juhani Pitkänen, Jarno Vanne
QoMEX5
2023 Vectorized and Optimized Dependent Quantization for Practical VVC Encoding
abstract
Versatile Video Coding (VVC) introduces dependent quantization (DQ) as a key coding tool, but it also has a high computational complexity. Therefore, optimizing DQ is crucial for practical VVC encoder implementations. In this paper, we propose optimization techniques for reducing the complexity of the DQ process, including efficient structure initialization, vectorized state update, rate distortion (RD) cost calculation, and last nonzero (LNZ) coefficient identification. The proposed optimizations are incorporated into the uvg266 practical VVC encoder. According to our evaluations, these optimization techniques accelerate the DQ process by over 1.7× and the entire VVC intra encoder by 1.4×, without any coding efficiency degradation. The proposed optimizations achieve more than twice the speedup over the existing techniques.
Joose Sainio, Alexandre Mercat, Jarno Vanne
VCIP3
2023 Machine Learning Based Efficient QT-MTT Partitioning Scheme for VVC Intra Encoders
abstract
The next-generation Versatile Video Coding (VVC) standard introduces a new Multi-Type Tree (MTT) block partitioning structure that supports Binary-Tree (BT) and Ternary-Tree (TT) splits in both vertical and horizontal directions. This new approach leads to five possible splits at each block depth. It thereby improves the coding efficiency of VVC over that of the preceding High Efficiency Video Coding (HEVC) standard, which only supports Quad-Tree (QT) partitioning with a single split per block depth. However, MTT also has brought a considerable impact on encoder computational complexity. This paper proposes a two-stage learning-based technique to tackle the complexity overhead of MTT in VVC intra encoders. In our scheme, the input block is first processed by a Convolutional Neural Network (CNN) to predict its spatial features through a vector of probabilities describing the partition at each$4\times 4$edge. Subsequently, a Decision Tree (DT) model leverages this vector of spatial features to predict the most likely splits at each block. Finally, based on this prediction, only the$N$most likely splits are processed by the Rate-Distortion (RD) process of the encoder. In order to train our CNN and DT models on a wide range of image contents, we also propose a public VVC frame partitioning dataset based on existing image dataset encoded with the VVC reference software encoder. Our solution relying on the top-3 configuration reaches 47.4% complexity reduction for a negligible bitrate increase of 0.79%. A top-2 configuration enables a higher complexity reduction of 70.4% for 2.49% bitrate loss. These results emphasize a better trade-off between VTM intra-coding efficiency and complexity reduction compared to the state-of-the-art solutions. The source code of the proposed method and the training dataset are made publicly available at GitHub.
Alexandre Tissier, Wassim Hamidouche, Souhaiel Belhadj Dit Mdalsi, Jarno Vanne, Franck Galpin, Daniel Ménard
IEEE Trans. Circuits Syst. Video Technol.4
2022 Spatio-Temporal Parallelization Scheme for HEVC Encoding on Multi-Computer Systems
abstract
High Efficiency Video Coding (HEVC) sets the scene for economic video transmission and storage, but its inherent computational complexity calls for efficient parallelization techniques. This paper introduces and compares three different parallelization strategies for HEVC encoding on multi-computer systems: 1) spatial parallelization scheme, where input video frames are divided into slices and distributed among available computers; 2) temporal parallelization scheme, where input video is distributed among computers in groups of consecutive frames; 3) spatio-temporal parallelization scheme that combines the proposed spatial and temporal approaches. All these three schemes were benchmarked as part of the practical Kvazaar open-source HEVC encoder. Our experimental results on 2–5 computer configurations show that using the spatial scheme gives 1.65×–2.90× speedup at the cost of 4.16%–13.09% bitrate loss over a single-computer setup. The respective speedup with temporal parallelization is 1.86×–3.26× without any coding overhead. The spatio-temporal scheme with 2 slices was shown to offer the best load-balancing with 1.81×–3.55× speedups and a constant coding loss of 4.16%.
Alexandre Mercat, Sami Ahovainio, Jarno Vanne
ICIP3
2022 Machine Learning Based Efficient Qt-Mtt Partitioning for VVC Inter Coding
abstract
The Joint Video Experts Team (JVET) have standardized the Versatile Video Coding (VVC) in 2020 targeting efficient coding of the emerging video services and formats such as 8K and immersive video streaming applications. VVC standard enhances the coding efficiency by 40% at the cost of an encoder computational complexity increase estimated to 859%(x8) compared to the previous standard High Efficiency Video Coding (HEVC). This work aims at reducing the complexity of the VVC encoder under the Random Access (RA) configuration. The proposed method takes advantage of the inter prediction in order to predict the split probabilities through a convolutional neural network. Our solution reaches 31.8% of complexity reduction for a negligible bitrate increase of 1.11% outperforming state-of-the-art methods.
Alexandre Tissier, Wassim Hamidouche, Jarno Vanne, Daniel Ménard
ICIP3
2022 High-Level Synthesis Implementation of an Embedded Real-Time HEVC Intra Encoder on FPGA for Media Applications
abstract
High Efficiency Video Coding (HEVC) is the key enabling technology for numerous modern media applications. Overcoming its computational complexity and customizing its rich features for real-time HEVC encoder implementations, calls for automated design methodologies. This article introduces the first complete High-Level Synthesis (HLS) implementation for HEVC intra encoder on FPGA. The C source code of our open-source Kvazaar HEVC encoder is used as a design entry point for HLS that is applied throughout the whole encoder design process, from data-intensive coding tools like intra prediction and discrete transforms to more control-oriented tools such as context-adaptive binary arithmetic coding (CABAC). Our prototype is run on Nokia AirFrame Cloud Server equipped with 2.4 GHz dual 14-core Intel Xeon processors and two Intel Arria 10 PCIe FPGA accelerator cards with 40 Gigabit Ethernet. This proof-of-concept system is designed for hardware-accelerated HEVC encoding and it achieves real-time 4K coding speed up to 120 fps. The coding performance can be easily scaled up by adding practically any number of network-connected FPGA cards to the system. These results indicate that our HLS proposal not only boosts development time, but also provides previously unseen design scalability with competitive performance over the existing FPGA and ASIC encoder implementations.
Panu Sjövall, Ari Lemmetti, Jarno Vanne, Sakari Lahti, Timo Hämäläinen 0001
ACM Trans. Design Autom. Electr. Syst.3
2021 Parallel Implementations of Lambda Domain and R-Lambda Model Rate Control Schemes in a Practical HEVC Encoder
abstract
This paper summarizes the key parallelization aspects of the two most popular academic rate control (RC) algorithms for High Efficiency Video Coding (HEVC): 1) Lambda domain (LD) [1] and 2) R-Lambda model (R-LM) [2] that were originally designed for the HEVC test model (HM). In this work, these two sequential RC algorithms were ported to the practical Kvazaar open-source HEVC encoder [3] and adapted to work with wavefront parallel processing (WPP) and overlapping wavefront (OWF) parallelization techniques.
Joose Sainio, Alexandre Mercat, Jarno Vanne
DCC3
2021 Programmable Systems for Intelligence in Automobiles (PRYSTINE): Final results after Year 3
abstract
Autonomous driving is disrupting the automotive industry as we know it today. For this, fail-operational behavior is essential in the sense, plan, and act stages of the automation chain in order to handle safety-critical situations on its own, which currently is not reached with state-of-the-art approaches.The European ECSEL research project PRYSTINE realizes Fail-operational Urban Surround perceptION (FUSION) based on robust Radar and LiDAR sensor fusion and control functions in order to enable safe automated driving in urban and rural environments. This paper showcases some of the key exploitable results (e.g., novel Radar sensors, innovative embedded control and E/E architectures, pioneering sensor fusion approaches, AI-controlled vehicle demonstrators) achieved until its final year 3.
Norbert Druml, Anna Ryabokon, Rupert Schorn, Jochen Koszescha, Kaspars Ozols, Aleksandrs Levinskis, Rihards Novickis, Ethiopia Nigussie, Jouni Isoaho, Selim Solmaz, Georg Stettinger, Sergio E. Diaz, Mauricio Marcano, Jorge Villagra, Juan Medina, Martina Schwarz, Antonio Artuñedo, Mauro Comi, Rutger Beekelaar, Onur Özçelik, Elif Aksu Tasdelen, Yesim Gürbüz, Jan Saijets, Jukka Kyynäräinen, Dmitry Morits, Björn Debaillie, Maxim Rykunov, Joan Escamilla, Jarno Vanne, Tomi Korhonen, Kalle Holma, Eva-Maria Matzhold, Carlo Novara, Fabio Tango, Paolo Burgio, Giuseppe Carlo Calafiore, Milad Karimshoushtari, Emilie Boulay, Miguel Dhaens, Kylian Praet, Han Zwijnenberg, Henri Palm, David Aledo Ortega, Ercan Kalali, Tuomas Pensala, Arto Kyytinen, Morten Larsen, Omar Veledar, Georg Macher, Michael Lafer, Lorenzo Giraudi, Jakob Reckenzaun, Daniel Hammer, Naveen Mohan, Josef Schmid, Alfred Höß, Shai Ophir, Anand Dubey, Jonas Fuchs, Maximilian Lübke, Andrei Anghel, Nicolae-Catalin Ristea, Martin Törngren, Alua Musralina, Marlene Harter, Joseena Memadathil Jose, George Dimitrakopoulos 0001
DSD29
2021 High-Level Synthesis Implementation of Transform-Exempted SATD Architectures for Low-Power Video Coding
abstract
This paper presents the first known high-level synthesis (HLS) implementation for the Sum of Absolute Transformed Differences (SATD) calculation. The proposed hardware architecture is designed for two SATD algorithms: a widespread Fast Walsh-Hadamard Transform (FWHT-SATD) and a recently introduced Transform Exempted scheme (TE- SATD). This 2-stage architecture is made up of two 1-D Walsh- Hadamard Transform (WHT) stages and a transpose buffer (TB) between them. The chosen HLS approach cuts down design time over contemporary design methods and thereby made it feasible to implement a set of dedicated FWHT-SATD and TE- SATD architectures for 4x4, 8x8, and 16x16 pixel blocks. All these six architectures were synthesized for 28 nm and 45 nm standard cell technologies, and their area and energy consumptions were analysed. TE-based implementations provide 6.0-8.3% total cell area savings and 6.9-12.7% better energy-efficiency than traditional FWHT approaches. Our proposal is the first to introduce TE-SATD architectures for up to 16x16 blocks and each of these tailored architectures was shown to provide better trade-off between silicon area and performance than their reference implementations.
Tero Partanen, Ari Lemmetti, Panu Sjövall, Jarno Vanne
ISCAS4
2021 Open-source RTP Library for End-to-End Encrypted Real-Time Video Streaming Applications
abstract
Information security has become a key success factor for streaming media applications that are increasingly vulnerable to wiretapping, message forgery, data tampering, hacking, and other possible cyberattacks. This paper addresses the existing security risks in real-time video streaming by introducing a new security extension to our uvgRTP opensource Real-time Transport Protocol (RTP) library. The proposed solution improves content integrity and privacy by adopting Secure RTP (SRTP) and Zimmermann RTP (ZRTP) for media End-to-End Encryption (E2EE). These new security mechanisms make uvgRTP the first open-source library that supports on-the-fly encrypted AVC, HEVC, and VVC video streaming. Our performance results on Intel Core i7-4770 processor show that uvgRTP is able to transport encrypted 8K VVC video at up to 187 fps and 8K HEVC video at up to 120 fps over a 10 Gbps Local Area Network (LAN). The achieved transfer rate for encrypted HEVC video is 50% higher and latency 86% lower than the respective performance values of FFmpeg in unencrypted HEVC streaming. These top streaming speed results with state-of-the-art video codec support, advanced encryption mechanisms, and the permissive BSD license make uvgRTP an attractive solution for a broad range of commercial and academic streaming media applications.
Joni Räsänen, Aaro Altonen, Alexandre Mercat, Jarno Vanne
ISM4
2021 Open3DGen: open-source software for reconstructing textured 3D models from RGB-D images
abstract
This paper presents the first entirely open-source and cross-platform software called Open3DGen for reconstructing photorealistic textured 3D models from RGB-D images. The proposed software pipeline consists of nine main stages: 1) RGB-D acquisition; 2) 2D feature extraction; 3) camera pose estimation; 4) point cloud generation; 5) coarse mesh reconstruction; 6) optional loop closure; 7) fine mesh reconstruction; 8) UV unwrapping; and 9) texture projection. This end-to-end scheme combines multiple state-of-the-art techniques and provides an easy-to-use software package for real-time 3D model reconstruction and offline texture mapping. The main innovation lies in various Structure-from-Motion (SfM) techniques that are used with additional depth data to yield high-quality 3D models in real-time and at low cost. The functionality of Open3DGen has been validated on AMD Ryzen 3900X CPU and Nvidia GTX1080 GPU. This proof-of-concept setup attains an average processing speed of 15 fps for 720p (1280x720) RGBD input without the offline backend. Our solution is shown to provide competitive 3D mesh quality and execution performance with the state-of-the-art commercial and academic solutions.
Teo Niemirepo, Marko Viitanen, Jarno Vanne
MMSys3
2021 uvgVenctester: Open-Source Test Automation Framework for Comprehensive Video Encoder Benchmarking
abstract
The agile and efficient development of modern video encoders calls for automated testing methodologies. This paper presents the first-of-its-kind open-source test automation framework called uvgVenctester (github.com/ultravideo/uvgVenctester) that is designed for comprehensive performance and conformance testing of video encoders with the desired set of test video sequences. Our framework comes with built-in support for the popular AVC, HEVC, VVC, VP9, and AV1 video coding formats and the state-of-the-art HM, Kvazaar, x265, VTM, VVenC, SVT-VP9, and SVT-AV1 video encoders. Furthermore, there are no technical limitations of adopting other formats or encoders. The developers can evaluate the encoder of interest under the three primary usage scenarios: 1) conformance testing of the encoded bitstream; 2) rate-distortion-complexity comparison with the other encoders; and 3) systematic exploration of encoding parameters. The framework provides commonly used analysis tools to quantify encoding quality, speed, and bitrate with versatile set of absolute and comparative results such as Bjøntegaard Delta (BD)-Rate for PSNR, SSIM, and VMAF quality metrics. The supported output formats include CSV, graph, and comparison table. They ensure that the results are available in human and machine-readable formats. To the best of our knowledge, the proposed framework is currently the most comprehensive and modular open-source software toolset for video encoder benchmarking.
Joose Sainio, Alexandre Mercat, Jarno Vanne
MMSys3
2020 Programmable Systems for Intelligence in Automobiles (PRYSTINE): Technical Progress after Year 2
abstract
Autonomous driving has the potential to disruptively change the automotive industry as we know it today. For this, fail-operational behavior is essential in the sense, plan, and act stages of the automation chain in order to handle safety-critical situations by its own, which currently is not reached with state-of-the-art approaches.The European ECSEL research project PRYSTINE realizes Fail-operational Urban Surround perceptION (FUSION) based on robust Radar and LiDAR sensor fusion and control functions in order to enable safe automated driving in urban and rural environments. This paper showcases some of the key results (e.g., novel Radar sensors, innovative embedded control and E/E architectures, pioneering sensor fusion approaches, AI controlled vehicle demonstrators) achieved until year 2.
Norbert Druml, Björn Debaillie, Andrei Anghel, Nicolae-Catalin Ristea, Jonas Fuchs, Anand Dubey, Torsten Reissland, Maike Hartstem, Viktor Rack, Anna Ryabokon, Kaspars Ozols, Rihards Novickis, Aleksandrs Levinskis, Omar Veledar, Georg Macher, Johannes Jany-Luig, Selim Solmaz, Jakob Reckenzaun, Naveen Mohan, Shai Ophir, Georg Stettinger, Sergio E. Diaz, Mauricio Marcano, Jorge Villagra, Andrea Castellano, Rutger Beekelaar, Fabio Tango, Jarno Vanne, Kalle Holma, Oguz Icoglu, George Dimitrakopoulos 0001
DSD28
2020 Real-Time Implementation Of Scalable Hevc Encoder
abstract
This paper presents the first known open-source Scalable HEVC (SHVC) encoder for real-time applications. Our proposal is built on top of Kvazaar HEVC encoder by extending its functionality with spatial and signal-to-noise ratio (SNR) scalable coding schemes. These two scalability schemes have been optimized for real-time coding by means of three parallelization techniques: 1) wavefront parallel processing (WPP); 2) overlapped wavefront (OWF); and 3) AVX2-optimized upsampling. On an 8-core Xeon W-2145 processor, the proposed spatially scalable Kvazaar can encode twolayer 1080p video above 50 fps with scaling ratios of l.5 and 2. The respective coding gain s are 18.4% and 9.9% over Kvazaar simulcast coding at similar speed. Correspondingly, the coding speed of SNR scalable Kvazaar exceeds 30 fps with two-layer 1080p video. On average, it obtain s1.20 times speedup and 17.0% better coding efficiency over the simulcast case. These results justify the benefits of the proposed scalability schemes in real-time SHVC coding.
Jaakko Laitinen, Ari Lemmetti, Jarno Vanne
ICIP3
2020 CNN Oriented Complexity Reduction Of VVC Intra Encoder
abstract
The Joint Video Expert Team (JVET) is currently developing the next-generation MPEG/ITU video coding standard called Versatile Video Coding (VVC) and their ultimate goal is to double the coding efficiency over the state-of-the-art HEVC standard.The latest version of the VVC reference encoder, VTM6.1, is able to improve the intra coding efficiency by 24 % over the HEVC reference encoder HM16.20, but at the expense of 27 times the encoding time. The complexity overhead of VVC primarily stems from its novel block partitioning scheme that complements Quad-Tree (QT) split with Multi-Type Tree (MTT) partitioning in order to better fit the local variations of the video signal. This work reduces the block partitioning complexity of VTM6.1 through the use of Convolutional Neural Networks (CNNs). For each 64 × 64 Coding Unit (CU), the CNN is trained to predict a probability vector that speeds up coding block partitioning in encoding. Our solution is shown to decrease the intra encoding complexity of VTM6.1 by 51.5% with a bitrate increase of only 1.45%.
Alexandre Tissier, Wassim Hamidouche, Jarno Vanne, Franck Galpin, Daniel Ménard
ICIP3
2020 Live Demonstration: Multi-Laptop HEVC Encoding
abstract
This paper presents a demonstration setup for distributed real-time HEVC encoding on a multi-computer system. The demonstrated multi-level parallelization scheme is implemented in the practical Kvazaar open-source HEVC encoder. It allows Kvazaar to exploit parallelism at three levels: 1) Single Instruction Multiple Data (SIMD) optimized coding tools at the data level; 2) Wavefront Parallel Processing (WPP) and Overlapped Wavefront (OWF) parallelization strategies at the thread level; and 3) distributed slice encoding on multi-computer systems at the process level. This interactive demonstration allows visitors to gradually increase the degree of parallelism in Kvazaar and see the benefits of parallelization in live HEVC encoding. Exploiting all three levels of parallelism on a three-laptop setup speeds up Kvazaar by almost 21× over a non-parallelized single-core implementation of Kvazaar.
Sami Ahovainio, Alexandre Mercat, Jarno Vanne
ISCAS3
2020 Multi-Level Parallelization Scheme for Distributed HEVC Encoding on Multi-Computer Systems
abstract
High Efficiency Video Coding (HEVC) creates the conditions for cost-effective video transmission and storage but its inherent computational complexity calls for efficient parallelization techniques. This paper provides HEVC encoders with a holistic parallelization scheme that exploits parallelism at data, thread, and process levels at the same time. The proposed scheme is implemented in the practical Kvazaar open-source HEVC encoder. It makes Kvazaar exploit parallelism at three levels: 1) Single Instruction Multiple Data (SIMD) optimized coding tools at the data level; 2) Wavefront Parallel Processing (WPP) and Overlapped Wavefront (OWF) parallelization strategies at the thread level; and 3) distributed slice encoding on multi-computer systems at the process level. Our results show that the proposed process-level parallelization scheme increases the coding speed of Kvazaar by 1.86× on two computers and up to 3.92× on five computers with +0.19% and +0.81% coding losses, respectively. Exploiting all these three parallelism levels on a five-computer setup gives almost a 25× speedup over a non-parallelized single-core implementation.
Sami Ahovainio, Alexandre Mercat, Marko Viitanen, Jarno Vanne
ISCAS4
2020 Live Demonstration: Interactive Quality of Experience Evaluation in Kvazzup Video Call
abstract
This paper presents an interactive demonstration setup, which allows users to configure the video coding parameters of Kvazzup open-source video call software at runtime and evaluate their impact on Quality of Service (QoS) and Quality of Experience (QoE). The demonstration is carried out by implementing a new Kvazzup control panel for video call parameterization and visual quality, bit rate, latency, and frame rate evaluation.
Joni Räsänen, Aaro Altonen, Alexandre Mercat, Jarno Vanne
ISM4
2020 Binocular Multi-CNN System for Real-Time 3D Pose Estimation
abstract
The current practical approaches for depth-aware pose estimation convert a human pose from a monocular 2D image into 3D space with a single computationally intensive convolutional neural network (CNN). This paper introduces the first open-source algorithm for binocular 3D pose estimation. It uses two separate lightweight CNNs to estimate disparity/depth information from a stereoscopic camera input. This multi-CNN fusion scheme makes it possible to perform full-depth sensing in real time on a consumer-grade laptop even if parts of the human body are invisible or occluded. Our real-time system is validated with a proof-of-concept demonstrator that is composed of two Logitech C930e webcams and a laptop equipped with Nvidia GTX1650 MaxQ GPU and Intel i7-9750H CPU. The demonstrator is able to process the input camera feeds at 30 fps and the output can be visually analyzed with a dedicated 3D pose visualizer.
Teo Niemirepo, Marko Viitanen, Jarno Vanne
ACM Multimedia3
2020 Open-Source RTP Library for High-Speed 4K HEVC Video Streaming
abstract
Efficient transport technologies for High Efficiency Video Coding (HEVC) are key enablers for economic 4K video transmission in current telecommunication networks. This paper introduces a novel open-source Real-time Transport Protocol (RTP) library called uvgRTP for high-speed 4K HEVC video streaming. Our library supports the latest RFC 3550 specification for RTP and an associated RFC 7798 RTP payload format for HEVC. It is written in C++ under a permissive 2-clause BSD license and it can be run on both Linux and Windows operating systems with a user-friendly interface. Our experiments on an Intel Core i7-4770 CPU show that uvgRTP is able to stream HEVC video at 5.0 Gb/s over a local 10 Gb/s network. It attains 4.4 times as high peak goodput and 92.1% lower latency than the state-of-the-art FFmpeg multimedia framework. It also outperforms LIVE555 with over double the goodput and 82.3% lower latency. These results indicate that uvgRTP is currently the fastest open-source RTP library for 4K HEVC video streaming.
Aaro Altonen, Joni Räsänen, Jaakko Laitinen, Marko Viitanen, Jarno Vanne
MMSP5
2020 Kvazaar 2.0: fast and efficient open-source HEVC inter encoder
abstract
High Efficiency Video Coding (HEVC) is the key to economic video transmission and storage in the current multimedia applications but tackling its inherent computational complexity requires powerful video codec implementations. This paper presents Kvazaar 2.0 HEVC encoder that is the new release of our academic open-source software (github.com/ultravideo/kvazaar). Kvazaar 2.0 introduces novel inter coding functionality that is built on advanced rate-distortion optimization (RDO) scheme and speeded up with several early termination mechanisms, SIMD-optimized coding tools, and parallelization strategies. Our experimental results show that the proposed coding scheme makes Kvazaar 125 times as fast as the HEVC reference software HM on the Intel Xeon E5-2699 v4 22-core processor at the additional coding cost of only 2.4% on average. In constant quantization parameter (QP) coding, Kvazaar is also 3 times as fast as the respective preset of the well-known practical x265 HEVC encoder and is still able to attain 10.7% lower average bit rate than x265 for the same objective visual quality. These results indicate that Kvazaar has become one of the leading open-source HEVC encoders in practical high-efficiency video coding.
Ari Lemmetti, Marko Viitanen, Alexandre Mercat, Jarno Vanne
MMSys4
2020 UVG dataset: 50/120fps 4K sequences for video codec analysis and development
abstract
This paper provides an overview of our open Ultra Video Group (UVG) dataset that is composed of 16 versatile 4K (3840×2160) test video sequences. These natural sequences were captured either at 50 or 120 frames per second (fps) and stored online in raw 8-bit and 10-bit 4:2:0 YUV formats. The dataset is published on our website (ultravideo.cs.tut.fi) under a non-commercial Creative Commons BY-NC license. In this paper, all UVG sequences are described in detail and characterized by their spatial and temporal perceptual information, rate-distortion behavior, and coding complexity with the latest HEVC/H.265 and VVC/H.266 reference video codecs. The proposed dataset is the first to provide complementary 4K sequences up to 120 fps and is therefore particularly valuable for cutting-edge multimedia applications. Our evaluations also show that it comprehensively complements the existing 4K test set in VVC standardization, so we recommend including it in subjective and objective quality assessments of next-generation VVC codecs.
Alexandre Mercat, Marko Viitanen, Jarno Vanne
MMSys3
2019 Acceleration of Kvazaar HEVC Intra Encoder With Machine Learning
abstract
The complexity of High Efficiency Video Coding (HEVC) poses a real challenge to HEVC encoder implementations. Particularly, the complexity stems from the HEVC quad-tree structure that also has an integral part in HEVC coding efficiency. This paper presents a Machine Learning (ML) based technique for pruning the HEVC quad-tree without deteriorating coding gain. We show how ML decision trees can be used to predict a depth interval for a quad-tree before the Rate-Distortion Optimization (RDO). This approach limits the number of RDO candidates and thus speeds up encoding. The proposed technique works particularly well with high-quality video coding and it is shown to accelerate the veryslow preset of practical Kvazaar HEVC intra encoder by 1.35× with 0.49% bit rate increase. Compared with the corresponding preset of x265 encoder, Kvazaar is 2.12× as fast at a cost of under 1.21% bit rate overhead. These results indicate that the optimized Kvazaar is the leading open-source encoder in high-quality HEVC intra coding.
Alexandre Mercat, Ari Lemmetti, Marko Viitanen, Jarno Vanne
ICIP4
2019 Remote VR Gaming on Mobile Devices
abstract
This paper presents a remote 360-degree virtual reality (VR) gaming system for mobile devices. In this end-to-end scheme, execution of VR game is off-loaded from low-power mobile devices to a remote server where the executed game is rendered based on controller orientation and actions transmitted over the network. The server is running the Unity game engine and Kvazaar video encoder. Kvazaar compresses the rendered views of the game to High Efficiency Video Coding (HEVC) video that is streamed to a player over a regular WiFi link in real time. The frontend of our proof-of-concept demonstrator setup is composed of the Samsung Galaxy S8 smartphone and Google Daydream View VR headset with a controller. The backend server is a laptop equipped with Nvidia GTX 1070 GPU and Intel i7 7820HK CPU. The system is able to run the demonstrated 360-degree shooting VR game with 1080p resolution at 30 fps while keeping motion-to-photon latency close to 50 ms. This approach lets players enter immersive gaming experience without a need to invest in all-in-one VR headsets.
Mikko Juhani Pitkänen, Marko Viitanen, Alexandre Mercat, Jarno Vanne
ACM Multimedia4
2019 Complexity Reduction Opportunities in the Future VVC Intra Encoder
abstract
The Joint Video Expert Team (JVET) is developing the next-generation video coding standard called Versatile Video Coding (VVC) and their ultimate goal is to double the coding efficiency over the current state-of-the-art standard HEVC without letting complexity get out of hand. This work addresses the complexity of the VVC reference encoder called VVC Test Model (VTM) under All Intra coding configuration. The VTM3.0 is able to improve intra coding efficiency by 21% over the latest HEVC reference encoder HM16.19. This coding gain primarily stems from three new coding tools. First, the HEVC Quad-Tree (QT) structure extension with Multi-Type Tree (MTT) partitioning. Second, the duplication of intra prediction modes from 35 to 67. And third, the Multiple Transform Selection (MTS) scheme with two new discrete cosine/sine transforms (DCT-VIII and DST-VII). However, these new tools also play an integral part in making VTM intra encoding around 20 times as complex as that of HM. The purpose of this work is to analyze these tools individually and specify theoretical upper limits for their complexity reduction. According to our evaluations, the complexity reduction opportunity of block partitioning is up to 97%, i.e., the encoding complexity would drop down to 3% for the same coding efficiency if the optimal block partitioning could be directly predicted. The respective percentages for intra mode reduction and MTS optimization are 65% and 55%. We believe these results motivate VVC codec designers to develop techniques that are able to take most out of these opportunities.
Alexandre Tissier, Alexandre Mercat, Thomas Amestoy, Wassim Hamidouche, Jarno Vanne, Daniel Ménard
MMSP5
2019 Public and open HEVC encoding service in the cloud
abstract
The ability to record vast amounts of video content requires convenient and efficient video coding services with which users can tackle the limited storage and transmission capacities. This paper presents an open-source cloud service for encoding raw video formats and transcoding compressed videos to the latest HEVC/H.265 format. Respective commercial transcoding services are available on the Internet but they are behind a paywall. On the other hand, using command-line interfaces of existing open-source software solutions requires in-depth knowledge of the coding process to attain the best coding gain and speed. The proposed service is available online, it is free to use without any registration, and its easy-to-use web interface makes it feasible for non-technical users. It is built on the FFmpeg multimedia framework whose built-in decoders accept various input video formats that are then compressed to HEVC with a full-fledged Kvazaar open-source encoder.
Aaro Altonen, Marko Viitanen, Joni Räsänen, Alexandre Mercat, Jarno Vanne
MMSys5
2019 Parallax-Tolerant 360 Live Video Stitcher
abstract
This paper presents an open-source software implementation for real-time 360-degree video stitching. To ensure a seamless stitching result, cylindrical and content-preserving warping are implemented to dynamically correct image alignment and parallax, which may drift due to scene changes, moving objects, or camera movement. Depth variation, color changes, and lighting differences between adjacent frames are also smoothed out to improve visual quality of the panoramic video. The system is benchmarked with six 1080p videos, which are stitched into 4096×732 pixel output format. The proposed algorithm attains an output rate of 18 frames per second on GeForce GTX 1070 GPU and real-time speed can be met with a high-end GPU.
Miko Atokari, Marko Viitanen, Alexandre Mercat, Emil Kattainen, Jarno Vanne
VCIP5
2019 Visualization of Dynamic Resource Allocation for HEVC Encoding in FPGA-Accelerated SDN Cloud
abstract
This paper describes a demonstration setup to visualize dynamic resource allocation for real-time HEVC encoding services in FPGA-accelerated cloud. The demonstrated application is Kvazaar HEVC intra encoder, whose functionality is partitioned between FPGAs and processors. During the demonstration, several encoding services can be invoked with requests to the resource manager, which is responsible for allocation, deallocation, and load balancing of resources in the network. The manager provides JSON data to the visualizer, which uses D3 JavaScript library to visualize 1) the physical network structure; 2) running services; and 3) performance of the network elements. This interactive demonstration allows users to request new video streams, view the encoded streams, observe the visualization of the network and services, and manually turn on/off resources to test the robustness of the system.
Panu Sjövall, Mikko Teuho, Arto Oinonen, Jarno Vanne, Timo Hämäläinen 0001
VCIP4
2019 Are We There Yet? A Study on the State of High-Level Synthesis
abstract
To increase productivity in designing digital hardware components, high-level synthesis (HLS) is seen as the next step in raising the design abstraction level. However, the quality of results (QoRs) of HLS tools has tended to be behind those of manual register-transfer level (RTL) flows. In this paper, we survey the scientific literature published since 2010 about the QoR and productivity differences between the HLS and RTL design flows. Altogether, our survey spans 46 papers and 118 associated applications. Our results show that on average, the QoR of RTL flow is still better than that of the state-of-the-art HLS tools. However, the average development time with HLS tools is only a third of that of the RTL flow, and a designer obtains over four times as high productivity with HLS. Based on our findings, we also present a model case study to sum up the best practices in comparative studies between HLS and RTL. The outcome of our case study is also in line with the survey results, as using an HLS tool is seen to increase the productivity by a factor of six. In addition, to help close the QoR gap, we present a survey of literature focused on improving HLS. Our results let us conclude that HLS is currently a viable option for fast prototyping and for designs with short time to market.
Sakari Lahti, Panu Sjövall, Jarno Vanne, Timo Hämäläinen 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2018 Rate-Distortion-Complexity Optimized Coding Scheme for Kvazaar HEVC Intra Encoder
abstract
This paper summarizes a low-complexity rate-distortion optimization (RDO) scheme for Kvazaar HEVC intra encoder (github.com/ultravideo/kvazaar). Our work particularly addresses RDO quantization (RDOQ) since it is the most complex intra coding tool taking almost 60% of the Kvazaar complexity.
Ari Lemmetti, Eemeli Kallio, Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001
DCC4
2018 Low Latency Edge Rendering Scheme for Interactive 360 Degree Virtual Reality Gaming
abstract
This paper describes the core functionality and a proof-of-concept demonstration setup for remote 360 degree stereo virtual reality (VR) gaming. In this end-to-end scheme, the execution of a VR game is off-loaded from an end user device to a cloud edge server in which the executed game is rendered based on user's field of view (FoV) and control actions. Headset and controller feedback is transmitted over the network to the server from which the rendered views of the game are streamed to a user in real-time as encoded HEVC video frames. This approach saves energy and computation load of the end terminals by making use of the latest advancements in network connection speed and quality. In the showcased demonstration, a VR game is run in Unity on a laptop powered by i7 7820HK processor and GTX 1070 GPU. The 360 degree spherical view of the game is rendered and converted to a rectangular frame using equirectangular projection (ERP). The ERP video is sliced vertically and only the FoV is encoded with Kvazaar HEVC encoder in real time and sent over the network in UDP packets. Another laptop is used for playback with a HTC Vive VR headset. Our system can reach an end-to-end latency of 30 ms and bit rate of 20 Mbps for stereo 1080p30 format.
Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001, Ari Kulmala
ICDCS2
2018 Live Demonstration: End-to-End Real-Time ROI-based Encryption in HEVC Videos
abstract
This paper presents a demonstration setup for live HEVC video coding with Region of Interest (ROI) encryption. The showcased approach splits video frames into independent HEVC tiles and encrypts those belonging to the ROI. This end-to-end content protection scheme is put into practice by integrating the algorithms of selective encryption into Kvazaar HEVC encoder and decryption into openHEVC decoder. The shown implementation performs secure encryption of the ROI in real time with small bit rate and complexity overhead.
Naty Ould Sidaty, Marko Viitanen, Wassim Hamidouche, Jarno Vanne, Olivier Déforges
ISCAS4
2018 FPGA-Powered 4K120p HEVC Intra Encoder
abstract
This paper presents a hardware-accelerated Kvazaar HEVC intra encoder for 4K real-time video coding at up to 120 fps. The encoder is implemented on a Nokia AirFrame Cloud Server featuring a 2.4 GHz dual 14-core Intel Xeon processor and two Arria 10 PCI Express FPGA accelerator cards. The presented encoder is a speed-optimized version of our 1st generation 4K40p HEVC intra encoder. The proposed speedup techniques include 1) Increasing the number of FPGA cards to two; 2) Remapping the simplest multiplications from DSP blocks to logic for better FPGA utilization; 3) Making task scheduling more flexible to improve utilization rate of hardware accelerators; and 4) Increasing the pipeline depth and duplicating time-sensitive resources in the hardware accelerator. As a result, up to three hardware accelerator instances can be accommodated in a single Arria 10 so the encoder is able to make use of six accelerators. According to our experiments, the proposed encoder obtains threefold speedup over our 1st generation encoder. Our proposal is also shown to outperform all other encountered FPGA and ASIC implementations.
Panu Sjövall, Vili Viitamäki, Jarno Vanne, Timo Hämäläinen 0001, Ari Kulmala
ISCAS3
2018 Live Demonstration: 4K100p HEVC Intra Encoder
abstract
This paper describes a demonstration setup for real-time 4K HEVC intra coding. The system is built on Kvazaar open-source HEVC encoder partitioned between 22-core Xeon processor and two Arria 10 FPGAs. The demonstrator supports 1) live streaming of up to three 4K30p videos; or 2) offline video streaming up to 4K100p format. Live feeds are shot by three cameras whereas offline video is accessed from a local hard drive. In both cases, encoded bit stream is sent over a wired connection and played back by laptop(s). The demonstrated HEVC coding speed is over three times as fast as that of a pure software solution.
Vili Viitamäki, Panu Sjövall, Jarno Vanne, Timo Hämäläinen 0001, Ari Kulmala
ISCAS3
2018 Live Demonstration: Kvazzup 4K HEVC Video Call
abstract
This paper describes a demonstration setup for an end-to-end 4K video call with Kvazzup open-source HEVC video call application. The Kvazzup clients are installed on a desktop and a laptop computer powered by Intel 22-core Xeon and Intel 4-core i7 processors, respectively. The proposed two-way peer-to-peer video call setup is shown to support 2160p30 video stream from the desktop to the laptop and 720p30 stream in the reverse direction.
Joni Räsänen, Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001
ISM3
2018 Eye-Controlled Region of Interest HEVC Encoding
abstract
This paper presents a demonstrator setup for real-time HEVC encoding with gaze-based region of interest (ROI) detection. This proof-of-concept system is built on Kvazaar open-source HEVC encoder and Pupil eye tracking glasses. The gaze data is used to extract the ROI from live video and the ROI is encoded with higher quality than non-ROI regions. This demonstration illustrates that performing HEVC encoding with non-uniform quality reduces bit rate by 40-90% and complexity by 10-35% over that of the conventional approaches with negligible to minor deterioration in subjective quality.
Joose Sainio, Arttu Ylä-Outinen, Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001
ISM4
2018 Open framework for error-compensated gaze data collection with eye tracking glasses
abstract
Eye tracking is nowadays the primary method for collecting training data for neural networks in the Human Visual System modelling. Our recommendation is to collect eye tracking data from videos with eye tracking glasses that are more affordable and applicable to diverse test conditions than conventionally used screen based eye trackers. Eye tracking glasses are prone to moving during the gaze data collection but our experiments show that the observed displacement error accumulates fairly linearly and can be compensated automatically by the proposed framework. This paper describes how our framework can be used in practice with videos up to 4K resolution. The proposed framework and the data collected during our sample experiment are made publicly available.
Kari Siivonen, Joose Sainio, Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001
ISM4
2018 Fast and easy live video service setup using lightweight virtualization
abstract
The service broker provides service providers with virtualized services that can be initialized rapidly and scaled up or down on demand. This demonstration paper describes how a service provider can set up a new video distribution service to end users with a diminutive effort. Our proposal makes use of Docker lightweight virtualization technologies that pack services in containers. This makes it possible to implement video coding and content delivery networks that are scalable and consume resources only when needed. The demonstration showcases a scenario where a video service provider sets up a new live video distribution service to end users. After the setup, live 720p30 video camera feed is encoded in real-time, streamed in HEVC MPEG-DASH format over CDN network, and accessed with a HbbTV compatible set-top-box. This end-to-end system illustrates that virtualization causes no significant resource or performance overhead but is a perfect match for online video services.
Antti Heikkinen, Pekka Pääkkönen, Marko Viitanen, Jarno Vanne, Tommi Riikonen, Kagan Bakanoglu
MMSys4
2017 High-level synthesis implementation of HEVC 2-D DCT/DST on FPGA
abstract
This paper presents the first known high-level synthesis (HLS) implementation of integer discrete cosine transform (DCT) and discrete sine transform (DST) for High Efficiency Video Coding (HEVC). The proposed approach implements these 2-D transforms by two successive 1-D transforms using a well-known row-column and Even-Odd decomposition techniques. Altogether, the proposed architecture is composed of a 4-point DCT/DST unit for the smallest transform blocks (TBs), an 8/16/32-point DCT unit for the other TBs, and a transpose memory for intermediate results. On Arria II FPGA, the low-cost variant of the proposed architecture is able to support encoding of 1080p format at 60 fps and at the cost of 10.0 kALUTs and 216 DSP blocks. The respective figures for the proposed high-speed variant are 2160p at 30 fps with 13.9 kALUTs and 344 DSP blocks. These cost-performance characteristics outperform respective non-HLS approaches on FPGA.
Panu Sjövall, Vili Viitamäki, Jarno Vanne, Timo Hämäläinen 0001
ICASSP3
2017 High-level synthesized 2-D IDCT/IDST implementation for HEVC codecs on FPGA
abstract
This paper presents efficient inverse discrete cosine transform (IDCT) and inverse discrete sine transform (IDST) implementations for High Efficiency Video Coding (HEVC). The proposal makes use of high-level synthesis (HLS) to implement a complete HEVC 2-D IDCT/IDST architecture directly from the C code of a well-known Even-Odd decomposition algorithm. The final architecture includes a 4-point IDCT/IDST unit for the smallest transform blocks (TB), an 8/16/32-point IDCT unit for the other TBs, and a transpose memory for intermediate results. On Arria II FPGA, it supports real-time (60 fps) HEVC decoding of up to 2160p format with 12.4 kALUTs and 344 DSP blocks. Compared with the other existing HLS approach, the proposed solution is almost 5 times faster and is able to utilize available FPGA resources better.
Vili Viitamäki, Panu Sjövall, Jarno Vanne, Timo Hämäläinen 0001
ISCAS3
2017 Kvazzup: Open Software for HEVC Video Calls
abstract
This paper introduces an open-source HEVC video call application called Kvazzup. This academic proposal is the first HEVC-based end-to-end video call system with a user-friendly Graphical User Interface for call management. Kvazzup is built on the Qt framework and it makes use of four open-source tools: Kvazaar for HEVC encoding, OpenHEVC for HEVC decoding, Opus codec for audio coding, and Live555 for managing RTP/RTCP traffic. In our experiments, Kvazzup is prototyped with low-complexity VGA and high-quality 720p video calls between two desktops. On an Intel 4-core i5 processor, the VGA call accounts for 17% of the total CPU time. Averagely, it requires a bit rate of 0.31 Mbit/s out of which 0.26 Mbit/s is taken by video and 0.05 Mbit/s by audio. In the 720p call, the respective figures are 46%, 1.13 Mbit/s, 1.08 Mbit/s, and 0.05 Mbit/s. These test cases also validate the feasibility of HEVC in different types of video calls. HEVC coding is shown to account for around 34% of the Kvazzup processing time in the VGA call and 45% in the 720p call.
Joni Räsänen, Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001
ISM3
2017 Kvazaar: HEVC/H.265 4K30p Intra Encoder
abstract
This paper demonstrates the usage of Kvazaar open-source HEVC intra encoder in 4K real-time video encoding. In this setup, a raw 4K video is shot by an action camera, captured by an HDMI capture card, encoded in real-time by Kvazaar ultrafast preset on a 22-core Intel Xeon processor, sent to a laptop, and decoded by OpenHEVC decoder for playback. The encoding process is visualized on the fly by Kvazaar run-time visualizer.
Arttu Ylä-Outinen, Ari Lemmetti, Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001
ISM4
2016 AVX2-optimized Kvazaar HEVC intra encoder
abstract
This paper presents efficient SIMD optimizations for the open-source Kvazaar HEVC intra encoder. The C implementation of Kvazaar is accelerated by Intel AVX2 instructions whose effect on Kvazaar ultrafast preset is profiled. According to our profiling results, C functions of SATD, DCT, quantization, and intra prediction account for over 60% of the total intra coding time of Kvazaar ultrafast preset. This work shows that optimizing primarily these functions doubles the coding speed of a single-threaded Kvazaar intra encoder for the same rate-distortion performance. The highest performance boost is obtained by deploying the proposed optimizations jointly with multithreading. On the Intel 8-core i7 processor, the AVX2-optimized 16-threaded Kvazaar ultrafast preset achieves real-time (30 fps) intra coding speed up to 1080p resolution. Compared to AVX2-optimized ultrafast preset of x265, Kvazaar is 20% times faster and still obtains 9.1% bit rate gain for the same quality. These results justify that Kvazaar is currently the leading open-source HEVC intra encoder in terms of real-time coding speed and efficiency.
Ari Lemmetti, Ari Koivula, Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001
ICIP4
2016 Designing a clock cycle accurate application with high-level synthesis
abstract
During the recent years, high-level synthesis (HLS) has gained traction as a viable alternative to traditional handwritten register transfer level code in describing digital systems. This has been attributed to the maturing of the HLS tools and improving quality of their results. However, most published applications are data path intensive as HLS offers good tools for loop optimization, such as pipelining and loop unrolling. HLS is seldom applied to control-oriented applications since clock is not explicitly present in HLS source code. In this paper, we show how a clock cycle accurate application can be described with HLS. We give as a proof of concept an implementation of an FPGA-based I2C bus controller for an audio codec using Catapult C, and present a generalized work flow. Compared with a corresponding handwritten VHDL implementation, the HLS version consumes 84% more area at the same performance but productivity is increased by 100% at the first design time and even more with further design iterations.
Sakari Lahti, Jarno Vanne, Timo Hämäläinen 0001
IECON2
2016 Live demonstration: Run-time visualization of Kvazaar HEVC intra encoder
abstract
This demonstrator presents a run-time visualization tool for Kvazaar HEVC intra encoder. The implemented open-source tool is seamlessly integrated into Kvazaar to provide instant visual feedback of the encoding process. The visualization overlays the reconstructed HEVC video with the boundaries of the used block partitioning structure and associated intra prediction modes. The tool is also able to illustrate all Kvazaar parallelization schemes at run time: Wavefront Parallel Processing, tiles, and picture-level parallel processing. The displayed visualization information can be gradually adjusted to user needs. The tool is primarily designed for Kvazaar debugging but it also suits educational purposes.
Marko Viitanen, Ari Koivula, Jarno Vanne, Timo Hämäläinen 0001
ISCAS3
2016 RTP/RTCP Reception Hint Tracks for Video Call Recording and Playback
abstract
This paper proposes to use RTP (Real-time Transport Protocol) Reception Hint Tracks for convenient recording and playback of a video call in MP4 format. The feasibility of RTP Reception Hint Tracks is validated as a part of an implemented end-to-end video call system. The proposed approach records a bidirectional Linphone video call and multiplexes it as MP4 RTP Reception Hint Tracks with the L-SMASH software library. It also stores RTCP (RTP Control Protocol) Reception Hint Tracks for additional timing information. Playback of the MP4 file is performed with VLC Media Player that is made compatible with RTP Reception Hint Tracks. The proposed proof-of-concept setup meets particularly well the needs of multi-codec solutions where different audio and video codecs can be used for a video call recording and playback. According to our analysis, recording RTP reception Hint Tracks increases the Linphone CPU time by under 1% and the bitrate by under 2% over the bare bitrate of the recorded RTP packets.
Joni Räsänen, Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001, Miska M. Hannuksela, Vinod Kumar Malamal Vadakital
ISM3
2016 Kvazaar: Open-Source HEVC/H.265 Encoder
abstract
Kvazaar is an academic software video encoder for the emerging High Efficiency Video Coding (HEVC/H.265) standard. It provides students, academic professionals, and industry experts a free, cross-platform HEVC encoder for x86, x64, PowerPC, and ARM processors on Windows, Linux, and Mac. Kvazaar is being developed from scratch in C and optimized in Assembly under the LGPLv2.1 license. The development is being coordinated by Ultra Video Group at Tampere University of Technology (TUT) and the implementation work is carried out by an active community on GitHub. Developer friendly source code of Kvazaar makes joining easy for new developers. Currently, Kvazaar includes all essential coding tools of HEVC and its modular source code facilitates parallelization on multi and manycore processors as well as algorithm acceleration on hardware. Kvazaar is able to attain real-time HEVC coding speed up to 4K video on an Intel 14-core Xeon processor. Kvazaar is also supported by FFmpeg and Libav. These de-facto standard multimedia frameworks boost Kvazaar popularity and enable its joint usage with other well-known multimedia processing tools. Nowadays, Kvazaar is an integral part of teaching at TUT and it has got a key role in three Eureka Celtic-Plus projects in the fields of 4K TV broadcasting, virtual advertising, Video on Demand, and video surveillance.
Marko Viitanen, Ari Koivula, Ari Lemmetti, Arttu Ylä-Outinen, Jarno Vanne, Timo Hämäläinen 0001
ACM Multimedia5
2015 High-Level Synthesis Design Flow for HEVC Intra Encoder on SoC-FPGA
abstract
This paper presents a High-Level Synthesis (HLS) flow for mapping a software HEVC encoder into Altera CycloneV SoC-FPGA. The starting point is a C implementation of an open-source Kvazaar HEVC intra encoder, which is minimally refined for SystemC design space exploration and automatic Catapult-C RTL generation. The final implementation involves Kvazaar encoder executed in Linux on dual-core ARM, and HW accelerated intra prediction on FPGA. Changing the SW/HW partitioning or modifying the implementation takes hours instead of weeks with Catapult-C HLS. In addition, the design is portable to other platforms without major manual re-writing. We obtained 9 fps full-HD intra prediction speed with a single accelerator on Altera Cyclone V SX on Terasic VEEK-MT-C5SoC board including video capture and HEVC video streaming via Ethernet. To the best of our knowledge, this is the first reported HLS assisted implementation of HEVC encoder on SoC-FPGA.
Panu Sjövall, Janne Virtanen, Jarno Vanne, Timo Hämäläinen 0001
DSD3
2015 Kvazaar HEVC encoder for efficient intra coding
abstract
This paper presents an open-source Kvazaar encoder for HEVC intra coding. This academic software encoder has been developed from the scratch using C as an implementation language by prioritizing modularity, portability, and readability of the source code. Kvazaar implements almost the same intra coding functionality as HEVC reference encoder (HM) but its rewritten source code makes it significantly faster. In all-intra (AI) coding, a single-threaded C implementation of Kvazaar is 2.3 times faster than HM at a cost of 1.7% bit rate increase. The respective values with a high speed preset of Kvazaar are 10.6 and 8.8%. Compared to a single-threaded C++ implementation of x265, Kvazaar improves rate-distortion performance and increases encoding speed in both high-quality and high-speed test cases. Kvazaar has a particular edge in the high-speed test case where it almost halves the BD-rate loss and more than doubles the performance.
Marko Viitanen, Ari Koivula, Ari Lemmetti, Jarno Vanne, Timo Hämäläinen 0001
ISCAS4
2015 Kvazaar HEVC still image coding on Raspberry Pi 2 for low-cost remote surveillance
abstract
This demonstrator serves as a proof-of-concept of our multi-camera remote surveillance system that supports 1080p still image capture with a 10-second refresh rate. Image capture, compression, and broadcast are implemented in each Raspberry Pi 2 camera node. Image compression is conducted with an open-source Kvazaar HEVC encoder that outputs HEVC images in BPG format. The BPG images are broadcast from camera nodes to terminals over the Internet through WebSocket protocol. The images can be played back with most Web browsers in remote locations with Internet access.
Marko Viitanen, Ari Koivula, Jarno Vanne, Timo Hämäläinen 0001
VCIP3
2014 Comparative study of 8 and 10-bit HEVC encoders
abstract
This paper compares the rate-distortion-complexity (RDC) characteristics of the HEVC Main 10 Profile (M10P) and Main Profile (MP) encoders. The evaluations are performed with HEVC reference encoder (HM) whose M10P and MP are benchmarked with different resolutions, frame rates, and bit depths. The reported RD results are based on bit rate differences for equal PSNR whereas complexities have been profiled with Intel VTune on Intel Core 2 processor. With our 10-bit 4K 120 fps test set, the average bit rate decrements of M10P over MP are 5.8%, 11.6%, and 12.3% in the all-intra (AI), random access (RA), and low-delay B (LB) configurations, respectively. Decreasing the bit depth of this test set to 8 lowers the RD gain of Ml OP only slightly to 5.4% (AI), 11.4% (RA), and 12.1% (LB). The similar trend continues in all our tests even though the RD gain of M10P is decreased over MP with lower resolutions and frame rates. M10P introduces no computational overhead in HM, but it is anticipated to increase complexity and double the memory usage in practical encoders. Hence, the 10-bit HEVC encoding with 8-bit input video is the most recommended option if computation and memory resources are adequate for it.
Jarno Vanne, Marko Viitanen, Ari Koivula, Timo Hämäläinen 0001
VCIP1
2014 Efficient Mode Decision Schemes for HEVC Inter Prediction
abstract
The emerging High Efficiency Video Coding (HEVC) standard reduces the bit rate by almost 40% over the preceding state-of-the-art Advanced Video Coding (AVC) standard with the same objective quality but at about 40% encoding complexity overhead. The main reason for HEVC complexity is inter prediction that accounts for 60%-70% of the whole encoding time. This paper analyzes the rate-distortion-complexity characteristics of the HEVC inter prediction as a function of different block partition structures and puts the analysis results into practice by developing optimized mode decision schemes for the HEVC encoder. The HEVC inter prediction involves three different partition modes: square motion partition, symmetric motion partition (SMP), and asymmetric motion partition (AMP) out of which the decision of SMPs and AMPs are optimized in this paper. The key optimization techniques behind the proposed schemes are: 1) a conditional evaluation of the SMP modes; 2) range limitations primarily in the SMP sizes and secondarily in the AMP sizes; and 3) a selection of the SMP and AMP ranges as a function of the quantization parameter. These three techniques can be seamlessly incorporated in the existing control structures of the HEVC reference encoder without limiting its potential parallelization, hardware acceleration, or speed-up with other existing encoder optimizations. Our experiments show that the proposed schemes are able to cut the average complexity of the HEVC reference encoder by 31%-51% at a cost of 0.2%-1.3% bit rate increase under the random access coding configuration. The respective values under the low-delay B coding configuration are 32%-50% and 0.3%-1.3%.
Jarno Vanne, Marko Viitanen, Timo Hämäläinen 0001
IEEE Trans. Circuits Syst. Video Technol.1
2012 Complexity analysis of next-generation HEVC decoder
abstract
This paper analyzes the complexity of the HEVC video decoder being developed by the JCT-VC community. The HEVC reference decoder HM 3.1 is profiled with Intel VTune on Intel Core 2 Duo processor. The analysis covers both Low Complexity (LC) and High Efficiency (HE) settings for resolutions varying from WQVGA (416 × 240 pixels) up to 1600p (2560 × 1600 pixels). The yielded cycle-accurate results are compared with the respective results of H.264/AVC Baseline Profile (BP) and High Profile (HiP) reference decoders. HEVC offers significant improvement in compression efficiency over H.264/AVC: the average BD-rate saving of LC is around 51% over BP whereas the BD-rate gain of HE is around 45% over HiP. However, the average decoding complexities of LC and HE are increased by 61% and 87% over BP and HiP, respectively. In LC, the most complex functions are motion compensation (MC) and loop filtering (LF) that account on average for 50% and 14% of the decoder complexity. The decoding complexity of HE configuration is on average 42% higher than that of the LC configuration. Majority of the difference is caused by extra LF stages. In HE, the complexities of MC and LF are 37% and 32%, respectively. In practice, a standard 3 GHz dual core processor is expected to be able to decode 1080p HEVC content in real-time.
Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001, Moncef Gabbouj, Jani Lainema
ISCAS2
2012 Comparative Rate-Distortion-Complexity Analysis of HEVC and AVC Video Codecs
abstract
This paper analyzes the rate-distortion-complexity of High Efficiency Video Coding (HEVC) reference video codec (HM) and compares the results with AVC reference codec (JM). The examined software codecs are HM 6.0 using Main Profile (MP) and JM 18.0 using High Profile (HiP). These codes are benchmarked under the all-intra (AI), random access (RA), low-delayB(LB), and low-delayP(LP) coding configurations. In order to obtain a fair comparison, JM HiP anchor codec has been configured to conform to HM MP settings and coding configurations. The rate-distortion comparisons rely on objective quality assessments, i.e., bit rate differences for equal PSNR. The complexities of HM and JM have been profiled at the cycle level with Intel VTune on Intel Core 2 Duo processor. The coding efficiency of HEVC is drastically better than that of AVC. According to our experiments, the average bit rate decrements of HM MP over JM HiP are 23%, 35%, 40%, and 35% under the AI, RA, LB, and LP configurations, respectively. However, HM achieves its coding gain with a realistic overhead in complexity. Our profiling results show that the average software complexity ratios of HM MP and JM HiP encoders are 3.2× in the AI case, 1.2× in the RA case, 1.5× in the LB case, and 1.3× in the LP case. The respective ratios with HM MP and JM HiP decoders are 2.0×, 1.6×, 1.5×, and 1.4×. This paper also reveals the bottlenecks of HM codec and provides implementation guidelines for future real-time HEVC codecs.
Jarno Vanne, Marko Viitanen, Timo Hämäläinen 0001, Antti Hallapuro
IEEE Trans. Circuits Syst. Video Technol.1
2009 A Configurable Motion Estimation Architecture for Block-Matching Algorithms
abstract
This paper introduces a configurable motion estimation architecture for a wide range of fast block-matching algorithms (BMAs). Contemporary motion estimation architectures are either too rigid for multiple BMAs or the flexibility in them is implemented at the cost of reduced performance. The proposed architecture overcomes both of these limitations. The configurability of the proposed architecture is based on a new BMA framework that can be adjusted to support the desired set of BMAs. The chosen framework configuration is implemented by an intelligent control logic which is integrated to an efficient parallel memory system and distortion computation unit. The flexibility of the framework is demonstrated by mapping five different BMAs (BBGDS, DS, CDS, HEXBS, and TSS) to the architecture. The total execution time of the mapped BMAs is shown to be almost directly proportional to the number of tested checking points in the search area, so the architecture is very tolerant of different BMA-specific search strategies and search patterns. In addition, a run-time switching between supported BMAs can be done without performance compromises. With a 0.13-mum CMOS technology, the proposed architecture configured for HEXBS, BBGDS, and TSS requires only 14.2 kgates and 2.5 KB of memory at 200 MHz operating frequency. A performance comparison to the reference programmable architectures reveals that only the proposed implementation is able to process real-time (30 fps) fixed block-size motion estimation (1 reference frame) at full HDTV resolution (1920 times1080).
Jarno Vanne, Eero Aho, Kimmo Kuusilinna, Timo Hämäläinen 0001
IEEE Trans. Circuits Syst. Video Technol.1
2008 A Parallel Memory System for Variable Block-Size Motion Estimation Algorithms
abstract
This paper proposes an efficient parallel memory system for algorithms applied in fixed and variable block-size motion estimation (VBSME). The proposed system is implemented by a novel combination of two parallel memory architectures. The distribution of data among the memory modules is modified over contemporary approaches and the optimized address computation unit enables a rapid address generation for accessed memory locations. Furthermore, the introduced data permutation scheme organizes data efficiently for storage and retrieval. The proposed system enables up to 4 X speedup in data storage and retrieves data up to 55% faster for VBSME compared with the reference implementations. With a 0.18- mum CMOS technology, the proposed memory addressing and data permutation scheme can be clocked at 980 MHz operating frequency with a cost of less than 6 kgates. On FPGA, the system can operate at 200 MHz with less than 700 logic elements. The results show that the proposed system is applicable to real-time VBSME at HDTV resolution.
Jarno Vanne, Eero Aho, Timo Hämäläinen 0001, Kimmo Kuusilinna
IEEE Trans. Circuits Syst. Video Technol.1
2006 A High-Performance Sum of Absolute Difference Implementation for Motion Estimation
abstract
This paper presents a high-performance sum of absolute difference (SAD) architecture for motion estimation, which is the most time-consuming and compute-intensive part of video coding. The proposed architecture contains novel and efficient optimizations to overcome bottlenecks discovered in existing approaches. In addition, designed sophisticated control logic with multiple early termination mechanisms further enhance execution speed and make the architecture suitable for general-purpose usage. Hence, the proposed architecture is not restricted to a single block-matching algorithm in motion estimation, but a wide range of algorithms is supported. The proposed SAD architecture outperforms contemporary architectures in terms of execution speed and area efficiency. The proposed architecture with three pipeline stages, synthesized to a 0.18-mum CMOS technology, can attain 770-MHz operating frequency at a cost of less than 5600 gates. Correspondingly, performance metrics for the proposed low-latency 2-stage architecture are 730 MHz and 7500 gates
Jarno Vanne, Eero Aho, Timo Hämäläinen 0001, Kimmo Kuusilinna
IEEE Trans. Circuits Syst. Video Technol.1
2005 Comments on "Winscale: an image-scaling algorithm using an area pixel Model"
abstract
In the paper by Kim et al. (2003), the authors propose a new image scaling method called winscale. The presented method can be used for scaling up and down. However, scaling down utilizing the winscale concept gives exactly the same results as the well-known bilinear interpolation. Furthermore, compared to bilinear, scaling up with the proposed winscale "overlap stamping" method has very similar calculations. The basic winscale upscaling differs from the bilinear method.
Eero Aho, Jarno Vanne, Kimmo Kuusilinna, Timo Hämäläinen 0001
IEEE Trans. Circuits Syst. Video Technol.2
2002 Enhanced Configurable Parallel Memory Architecture
abstract
Contemporary multimedia processors and applications are increasingly limited by their data accessing capabilities. However, the designed Configurable Parallel Memory Architecture (CPMA) alleviates these multimedia data accessing requirements; achieving significant performance improvements over traditional memory architectures. CPMA decreases considerably the processor-memory bottleneck by widening the memory bandwidth, decreasing the number of memory accesses, and diminishing the significance of memory latency. To further enhance the performance of CPMA, this paper introduces a novel architectural extension called CPMA access instruction correlation recognition. The presented method is intended for accelerating the execution rate of consecutive, temporally conflict-free, CPMA memory accesses. As demonstrated in this paper, the superior CPMA performance can also be maintained in the case of limited access widths. In addition, the presented results confirm that CPMA can have an acceptable silicon area.
Jarno Vanne, Eero Aho, Kimmo Kuusilinna, Timo Hämäläinen 0001
DSD1