EDBT 2026 Demo / reviewers in the wild / expert
Alexandre Mercat
dblp:149/4675
· DBLP profile ↗
50ranked-venue papers
10as first author
28since 2021 · last 2026
0000-0003-2211-970XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 40 · 6 first-author · 26 since 2021Systems, architecture and hardware · 8 · 4 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UVG-VCM: Benchmarking Dataset for Machine-Oriented Visual Data Compression
Tero Partanen, Miro Anttila, Rudolf Kortelahti, Guillaume Gautier, Alexandre Mercat, Jarno Vanne |
QoMEX | 5 |
| 2026 | UVG-GS: Human-Centric Gaussian Splats Dataset for Visual Volumetric Compression
Uyen Phan, Alexandre Mercat, Patrice Rondao-Alface, Louis Fréneau, Jarno Vanne, Guillaume Gautier |
QoMEX | 2 |
| 2025 | Live Demonstration: VVC Intra-Coding VisualizerabstractThis paper presents a live demonstrator that is designed to visualize the encoding process of uvg266 VVC open-source intra encoder. The proposed visualizer provides real-time insights into the decision-making process of uvg266 by illustrating chosen block structures, modes, transforms, and parallelization strategies. The tool also facilitates interactive analysis through a user-friendly interface that allows users to dynamically adjust encoding tools and parameters on the fly. Researchers, developers, and students can use the tool to enhance their understanding of VVC coding flow and to support their optimization and educational efforts. Joose Sainio, Alexandre Mercat, Jarno Vanne |
ISCAS | 2 |
| 2025 | UVG-CWI-DQPC: Dual-Quality Point Cloud Dataset for Volumetric Video ApplicationsabstractVolumetric video is a key enabler of immersive extended reality (XR) experiences and is often represented using point clouds for their structural simplicity. However, capturing volumetric content through multi-view acquisition and depth sensing poses many challenges, such as occlusions and depth mismatches. To foster research in this field, we introduce a unique dual-quality point cloud dataset, named UVG-CWI-DQPC, which is designed to support the development of point cloud enhancement, compression, and quality assessment. Our dataset includes 12 dynamic sequences captured simultaneously by: 1) a high-end capture system producing high-fidelity point clouds with extensive processing; and 2) a consumer-grade capture system relying on affordable RGB-D cameras, lightweight processing, and open-source tools. For each sequence, our dataset provides ground-truth point clouds from the high-end capture system and raw RGB-D footage from the consumer-grade capture system, along with calibration data and tools for point cloud generation. This dual-quality setup enables direct comparison and benchmarking of algorithms for densification, occlusion removal, registration, and quality enhancement. Our dataset is publicly available under a permissive license to support reproducible research and standardization work in Moving Picture Experts Group (MPEG) and 3rd Generation Partnership Project (3GPP). Guillaume Gautier, Xuemei Zhou, Jack Jansen 0001, Louis Fréneau, Marko Viitanen, Uyen Phan, Jani Käpylä, Irene Viola 0001, Alexandre Mercat, Pablo César, Jarno Vanne |
ACM Multimedia | 10 |
| 2025 | uvgVPCCenc: Practical Open-Source Encoder for Fast V-PCC CompressionabstractVideo-based Point Cloud Compression (V-PCC) standard offers state-of-the-art tools and efficiency for volumetric video compression. The V-PCC reference software, TMC2, is able to demonstrate the compression efficiency of V-PCC, but with computational complexity that is impractical for real-life applications. This paper introduces a novel open-source V-PCC encoder, named uvgVPCCenc, that brings V-PCC compression speed to a practical level. To reduce computational bottlenecks, uvgVPCCenc is implemented from the ground up in C++ with advanced multi-threading and streamlined algorithms. It also supports a flexible parameterization and different presets that facilitate its adaptation to various content and target bitrates. Our comparative evaluations show that uvgVPCCenc achieves substantial speedups over TMC2 while maintaining competitive rate-distortion-complexity tradeoff. To the best of our knowledge, uvgVPCCenc is the first open-source encoder explicitly designed for practical V-PCC compression, paving the way from theoretical research to practical volumetric video applications. Louis Fréneau, Guillaume Gautier, Alexandre Mercat, Jarno Vanne |
MMSys | 3 |
| 2025 | Task Enhancement Tiles for Ultra Lightweight Post-processing in Visual Coding for MachinesabstractThe proliferation of automated visual analysis calls for compression methods tailored to the unique requirements of Video Coding for Machines (VCM). In this paper, we propose a computationally lightweight post-processing method that is based on a learned component referred to as a task enhancement tile (TET). A TET is spatially tiled over the reconstructed visual data and added to it element-wise. It only requires one addition per pixel in each color channel before the machine task can be applied. Our results with the VVC test model (VTM) demonstrate coding gains of up to 39.0% for object detection and 29.2% for instance segmentation on image datasets, while evaluation on a video dataset shows gains of up to 35.2% for object detection, relative to the VTM anchor. The proposed solution also offers extremely low computational cost, preservation of human-viewable content, full compliance with video coding standards, no requirement for side information transmission from encoder to decoder, and generalization across tasks, models, and encoding parameters. Tero Partanen, Alban Marie, Rudolf Kortelahti, Alexandre Mercat, Jarno Vanne, Miska M. Hannuksela, Honglei Zhang 0001, Alireza Aminlou, Francesco Cricri |
PCS | 4 |
| 2025 | Lightweight Motion-Vector-Based Segmentation Mask Tracking for Saliency-Based Rate ControlabstractIn video encoding, saliency-based rate control improves coding efficiency without compromising perceptual quality by allocating more bits to visually important regions. While object detection can guide bit allocation through rectangular bounding boxes, segmentation masks offer a more accurate delineation of object boundaries. This paper seeks to reduce the significant computational burden of frame-by-frame instance segmentation by proposing a lightweight segmentation mask tracking scheme, in which motion vectors (MVs) from the video encoder are used to predict per-vertex displacements. Altogether, we propose two neural network designs for segmentation mask tracking: (1) a base tracker optimized for tracking accuracy; and (2) a lite tracker that balances accuracy and computational complexity. Our experimental results show that the base tracker attains 70-88% of the accuracy of frame-by-frame instance segmentation but achieves a 48× speedup and reduces computational complexity to 0.03% on CPU. For the lite tracker, the corresponding figures are 49×, 0.01%, and 67-88%. Despite tradeoffs in tracking accuracy, reducing complexity to a fraction makes our solution a viable option for practical applications. Tero Partanen, Jarno Vanne, Alexandre Mercat |
VCIP | 4 |
| 2025 | High-Level Synthesis FPGA Implementation of Fractional Motion Estimation for HEVC and VVCabstractThe rapid deployment of modern video coding standards underscores the need for hardware implementations that enable real-time coding, support interoperability across codecs, and are openly accessible to community. This paper presents the first known high-level synthesis (HLS) implementation of an accurate full-search fractional motion estimation (FME). The proposed FME core is released as open-source and is compatible with High Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC) standards. It implements 1) an accurate multiplierless constant multiplication (MCM) unit that performs quarter-pixel interpolation over 9 × 9 pixels at a time; 2) a transform-exempted sum of absolute transformed differences (TE-SATD) unit that is optimized for area and speed; and 3) fixed-point Lagrangian optimizations for rate-distortion optimization (RDO). On an Intel Arria 10 FPGA, the FME core consumes 122 kALUTs and operates at up to 210 MHz. Our profiling results show that a single FME core can support practical HEVC and VVC encoding of 2160p video at 30–120 fps, depending on the video content and encoder preset. Jesse Smedberg, Panu Sjövall, Guillaume Gautier, Alexandre Mercat, Jarno Vanne |
VCIP | 4 |
| 2024 | Energy Reduction Opportunities in HDR Video EncodingabstractThis paper investigates the energy consumption of video encoding for high dynamic range videos. Specifically, we compare the energy consumption of the compression process using 10 -bit input sequences, a tone-mapped 8 -bit input sequence at 10 -bit internal bit depth, and encoding an 8 -bit input sequence using an encoder with an internal bit depth of 8 bit. We find that linear scaling of the luminance and chrominance values leads to degradations of the visual quality, but that significant encoding complexity and thus encoding energy can be saved. An important reason for this is the availability of vector instructions, which are not available for the 10-bit encoder. Furthermore, we find that at sufficiently low target bitrates, the compression efficiency at an internal bit depth of 8 bit exceeds the compression efficiency of regular 10-bit encoding. Christian Herglotz, Steven Le Moan, Alexandre Mercat |
ICIP | 3 |
| 2024 | Feasibility Study of Multi-Layer VVC Coding Scheme for Hybrid Machine-Human ConsumptionabstractThe proliferation of machine vision applications necessitates developing more efficient visual data compression schemes for machine consumption. However, numerous automated use cases still require keeping humans in the loop, leading to the need for a machine-optimized video streaming with the option for human supervision. This paper investigates the feasibility of using the multi-layer coding approach of the emerging Versatile Video Coding (VVC) standard to create favorable conditions for hybrid machine-human consumption. We introduce a multi-layer coding scheme, where the base layer (BL) is optimized for machines and the enhancement layer (EL) complements the stream for human vision. Our results demonstrate that the bitrate of the proposed multi-layer stream (BL + EL) is, on average, 11% higher than that of a single-layer VVC. However, the more compact BL yields overall bandwidth savings as long as the EL is required less than 80% of the time. Jaakko Laitinen, Tero Partanen, Alexandre Mercat, Jarno Vanne, Miska M. Hannuksela, Honglei Zhang 0001, Alireza Aminlou, Francesco Cricri |
ICME | 3 |
| 2024 | uvgComm: Open Software for Low-Latency Multi-party Video CommunicationabstractEffective and smooth video communication has become an integral tool for interaction in our modern society. The well-known WebRTC project is optimized for server-based communication architectures with less stringent latency requirements. This paper introduces an open-source, peer-to-peer (P2P) mesh-based application called uvgComm 1.0 for low-latency multi-party video communication. Our experimental results show that uvgComm attains 25% lower latency than server-based solutions at 219 ms. In addition, it offers advanced privacy protection mechanisms with end-to-end encryption and P2P mesh topology that delivers media without any media servers. The ambitious goal of uvgComm is to take P2P mesh-based conferencing architecture to the next level and make this architecture a serious competitor to server-based communication architectures. Joni Räsänen, Heikki Tampio, Alexandre Mercat, Jarno Vanne |
ACM Multimedia | 3 |
| 2024 | uvg266: Open-Source VVC Intra EncoderabstractVersatile Video Coding (VVC/H.266) standard is the emerging successor to the widespread High Efficiency Video Coding (HEVC/H.265). This work introduces the latest version of our academic open-source VVC intra encoder called uvg266. It has been developed from our well-known Kvazaar HEVC encoder by introducing new VVC coding tools into carefully optimized and parallelized coding flow of Kvazaar. This paper outlines the design methodology and implementation aspects of all intra (AI) configuration of uvg266. The experimental results show that single-threaded uvg266 is more than twice as fast as the state-of-the-art VVenC encoder in all our test cases. In speed-optimized coding, the coding overhead of uvg266 is 21.7% but the gap narrows down to 2.6% in the rate distortion optimized coding case. uvg266 has almost linear speedup with core count up to 32 threads. The better scalability of uvg266 quadruples the speed over VVenC. Furthermore, single-threaded uvg266 is up to 380× as fast as VVC reference software VTM and the gap raises to over 11,000× with 32 threads. To the best of our knowledge, uvg266 is currently the fastest available open-source VVC intra software encoder. Marko Viitanen, Joose Sainio, Kari Siivonen, Alexandre Mercat, Jarno Vanne |
ACM Multimedia | 4 |
| 2024 | Luma Range Scaling for Enhanced VVC Efficiency in Video Coding for MachinesabstractRecent years have shown significant growth in video data traffic for machine vision applications, catalyzing new standardization efforts in video coding for machines (VCM). These activities focus on compressing images and videos for machine vision tasks, rather than for human viewing. In this work, we propose a novel method that scales down the luma range to enhance the coding efficiency of Versatile Video Coding (VVC) for machine consumption. This method results in a lower bitrate after encoding and has only minimal adverse effects on the accuracy of machine vision tasks. In our experiments, we down-scale the luma channel of the input video using luma-scaling factors from 0.2 to 0.9 and evaluate coding results with optional back-scaling to the original range before machine vision tasks. Our results with the VVC Test Model (VTM) demonstrate that the proposed technique achieves coding gain of up to 37.9%and 46.1% for the same object detection and tracking accuracy, respectively. Tero Partanen, Alban Marie, Alexandre Mercat, Jarno Vanne, Miska M. Hannuksela, Honglei Zhang 0001, Alireza Aminlou, Francesco Cricri |
MMSP | 3 |
| 2024 | uvgRTP 3.0: Towards V3C Volumetric Video CommunicationabstractLow-latency volumetric video transport is a key enabling technology for more immersive communication applications. This paper presents the latest release of our open-source Real-time Transport Protocol (RTP) library called uvgRTP 3.0 that has been upgraded to support Visual Volumetric Video-based Coding (V3C) transmission in Video-based Point Cloud Compression (V-PCC) and MPEG Immersive Video (MIV) formats. uvgRTP 3.0 introduces 1) the V3C atlas RTP payload format; 2) two multiplexing methods to reduce port reservations during V3C transmission; and 3) improved packet reception with multithreading. Our performance results show that uvgRTP 3.0 can encrypt and transmit V3C bitstreams at 106 Mbit/s, with CPU core utilization of 26%, and achieving a round-trip latency of 2 ms in a local area network. The support for high-speed, encrypted V3C communication with the permissive BSD-license make uvgRTP 3.0 a potential transmission library for any industrial or academic volumetric communication system. Heikki Tampio, Joni Räsänen, Marko Viitanen, Alexandre Mercat, Guillaume Gautier, Jarno Vanne |
MMSys | 4 |
| 2024 | Tool Space Exploration of the V-PCC Patch Generation for Practical Point Cloud EncodingabstractVideo-based Point Cloud Compression (V-PCC) is the latest MPEG standard for visual volumetric data coding. V-PCC leverages the coding efficiency of well-established 2D video codecs by creating patches that link sub-regions of a point cloud to their 2D projections. This paper provides an in-depth analysis of the V-PCC patch generation process, which is the most compute-intensive part of the V-PCC standard. To the best of our knowledge, this is the first work to explore the patch generation tools individually and evaluate their impact on coding efficiency and complexity with different coding parameters. The main purpose of this characterization is to offer key insights into the V-PCC encoder design process and thereby foster the development of practical V-PCC encoders. Louis Fréneau, Alexandre Mercat, Guillaume Gautier, Joose Sainio, Jarno Vanne |
PCS | 2 |
| 2024 | Real-Time Video-based Point Cloud CompressionabstractThis paper presents a demonstration setup for our open-source intra encoder called uvgVPCCenc, which is optimized for real-time Video-based Point Cloud Compression (V-PCC). uvgVPCCenc achieves an average encoding speed of 26 frames per second (fps) on an Intel i7-12700 CPU when encoding volumetric video sequences with up to 185 000 points per frame. It is shown to be 700 times as fast as TMC2 reference implementation for V-PCC. Our work is the first to demonstrate real-time intra V-PCC encoding on a consumer-grade desktop computer. It indicates that even the immense computational complexity of intra V-PCC encoding can be tackled for practical applications with effective design and optimization techniques. Louis Fréneau, Guillaume Gautier, Heikki Tampio, Alexandre Mercat, Jarno Vanne |
VCIP | 4 |
| 2024 | Fast Machine Learning Aided Intra Mode Decision for Real-Time VVC Intra CodingabstractReducing the huge computational complexity of intra mode decision is the key to real-time Video Coding (VVC). This paper proposes a fast intra mode decision scheme that takes advantage of lightweight machine learning (ML) models to classify intra modes into fifteen clusters. The cluster is further refined using one of the three proposed strategies to select the most optimal mode. Our experimental results with the fastest configuration of the practical uvg266 encoder show that the proposed methods yield a competitive rate-distortion-complexity trade-off over a conventional rough mode decision (RMD). To the best of our knowledge, this is the first work to successfully reduce the complexity of RMD in a practical VVC encoder with the use of ML techniques. Joose Sainio, Baran Ataman, Alban Marie, Alexandre Mercat, Jarno Vanne |
VCIP | 4 |
| 2024 | Vectorized Angular Intra Prediction for Practical VVC EncodingabstractVersatile Video Coding (VVC) provides new coding tools for more efficient intra prediction but with a substantial increase in computational complexity. This paper introduces vectorized kernels for 8-bit angular intra prediction and position dependent intra prediction combination (PDPC), which are carefully optimized for all block sizes and prediction modes of VVC. The proposed kernels streamline the filtering process and utilize optimized memory access patterns. Our standalone tests show that the proposed vectorization achieves speedups of 6.68× for luma and 4.40× for chroma predictions over scalar implementations. Integrating these kernels into the practical uvg266 VVC encoder provides speedups of 1.07× in the slowest configuration and 1.68× in the fastest configuration. The reported speedups are obtained without any coding overhead, so the proposed vectorization plays an integral role in pursuing real-time VVC coding with high coding efficiency. Kari Siivonen, Joose Sainio, Guillaume Gautier, Alexandre Mercat, Jarno Vanne |
VCIP | 4 |
| 2023 | RDO Candidate Selection for Maximizing Coding Efficiency in a Practical HEVC EncoderabstractHigh Efficiency Video Coding (HEVC) creates the conditions for economic video transmission and storage but making most of its compression potential calls for effective rate-distortion optimization (RDO) techniques in practical HEVC encoders. This paper explores the effectiveness of the following universally applicable RDO techniques: 1) rough mode decision for intra RDO candidate selection; 2) number of intra and inter RDO search candidates; and 3) accurate bit cost estimation in entropy coding. All these techniques are implemented into Kvazaar open-source HEVC encoder and altogether they improve the coding efficiency of Kvazaar veryslow preset by 5.7%, 4.0%, and 7.1% with PSNR, SSIM, and VMAF quality metrics, respectively. Even though the proposed techniques reduce the coding speed of Kvazaar to 0.64×, Kvazaar is still, on average, 2.16× as fast as the x265 encoder and attains 12.4%, 22.3%, and 4.7% better coding gain for the same PSNR, SSIM, and VMAF quality, respectively. These results let us conclude that Kvazaar is currently the leading practical open-source solution for high-quality HEVC encoding. Joose Sainio, Alexandre Mercat, Jarno Vanne |
ICASSP | 2 |
| 2023 | FPGA-Accelerated HEVC Encoder for Energy-Efficient Multi-Access Edge ComputingabstractHigh Efficiency Video Coding (HEVC) and Multi-access Edge Computing (MEC) technologies can make real-time streaming media services available to users with reasonable bandwidth, but the computational complexity of HEVC tends to lead to increased energy consumption in these schemes. In this paper, we investigate the energy saving opportunities of utilizing a field-programmable gate array (FPGA) based HEVC encoder in edge media servers and devices. In practice, we analyze the energy impact of migrating our Kvazaar software HEVC intra encoder to Intel Arria 10 PCIe FPGA(s) on two platforms: 1) Nokia Airframe Cloud Server with 2.4 GHz dual 14-core Intel Xeon processors and 2) an embedded Jetson AGX Orin board with 2.2 GHz 12-core ARM processor. According to our experiments, FPGA encoding on these two platforms saved 76% and 86% of the energy taken up by software only encoding on Airframe, respectively. These results indicate the potential of FPGA-based video encoder acceleration in future green MEC architectures. Panu Sjövall, Alexandre Mercat, Jarno Vanne |
ICIP | 2 |
| 2023 | AVX2-Optimized Interpolation Filters for HEVC Inter EncodingabstractHigh Efficiency Video Coding (HEVC) sets the stage for economic video transmission and storage, but its inherent computational complexity calls for powerful implementations. This paper addresses the principal performance bottleneck of HEVC codecs by introducing AVX2-vectorized algorithms for HEVC interpolation filters. The proposed speed-up techniques include 1) a data permutation scheme for the horizontal interpolation stage; 2) a sliding window strategy for the vertical interpolation stage; 3) optimal usage of horizontal and vertical interpolation during fractional motion estimation; and 4) a lane-based approach to double the vector lengths from 128-bit legacy vector extensions to 256bits of AVX2. Our AVX2-optimized interpolation filters were benchmarked as part of the practical Kvazaar open-source HEVC encoder. On an Intel 8-core Xeon processor, they were shown to be 9.7 and 8.5 times as fast as scalar interpolation with the Kvazaar ultrafast and veryslow presets, respectively. In both cases, changing over from scalar to vectorized interpolation increases the coding speed of Kvazaar by more than 50%, which stresses the importance of interpolation optimizations in modern video encoders. Alexandre Mercat, Ari Lemmetti, Joose Sainio, Jarno Vanne |
ISCAS | 1 |
| 2023 | Open-Source Toolkit for Live End-to-End 4K VVC Intra CodingabstractVersatile Video Coding (VVC/H.266) takes video coding to the next level by doubling the coding efficiency over its predecessors for the same subjective quality, but at the cost of immense coding complexity. Therefore, VVC calls for aggressively optimized codecs to make it feasible for live streaming media applications. This paper introduces the first public end-to-end (E2E) pipeline for live 4K30p VVC intra coding and streaming. The pipeline is made up of three open-source components: 1) uvg266 for VVC encoding; 2) uvgRTP for VVC streaming; and 3) OpenVVC for VVC decoding. The proposed setup is demonstrated with a proof-of-concept prototype that implements the encoder end on AMD ThreadRipper 2990WX and the decoder end on Nvidia Jetson AGX Orin. Our prototype is almost 34 000 times as fast as the corresponding E2E pipeline built around the VTM codec. Respectively, it achieves 3.3 times speedup without any significant coding overhead over the pipeline that utilizes the fastest possible configuration of the well-known VVenC/VVdeC codec. These results indicate that our prototype is currently the only viable open-source solution for live 4K VVC intra coding and streaming. Marko Viitanen, Joose Sainio, Alexandre Mercat, Guillaume Gautier, Jarno Vanne, Ibrahim Farhat, Pierre-Loup Cabarat, Wassim Hamidouche, Daniel Ménard |
MMSys | 3 |
| 2023 | UVG-VPC: Voxelized Point Cloud Dataset for Visual Volumetric Video-based CodingabstractPoint cloud compression has become a crucial factor in immersive visual media processing and streaming. This paper presents a new open dataset called UVG-VPC for the development, evaluation, and validation of MPEG Visual Volumetric Video-based Coding (V3C) technology. The dataset is distributed under its own non-commercial license. It consists of 12 point cloud test video sequences of diverse characteristics with respect to the motion, RGB texture, 3D geometry, and surface occlusion of the points. Each sequence is 10 seconds long and comprises 250 frames captured at 25 frames per second. The sequences are voxelized with a geometry precision of 9 to 12 bits, and the voxel color attributes are represented as 8-bit RGB values. The dataset also includes associated normals that make it more suitable for evaluating point cloud compression solutions. The main objective of releasing the UVG-VPC dataset is to foster the development of V3C technologies and thereby shape the future in this field. Guillaume Gautier, Alexandre Mercat, Louis Fréneau, Mikko Juhani Pitkänen, Jarno Vanne |
QoMEX | 2 |
| 2023 | Vectorized and Optimized Dependent Quantization for Practical VVC EncodingabstractVersatile Video Coding (VVC) introduces dependent quantization (DQ) as a key coding tool, but it also has a high computational complexity. Therefore, optimizing DQ is crucial for practical VVC encoder implementations. In this paper, we propose optimization techniques for reducing the complexity of the DQ process, including efficient structure initialization, vectorized state update, rate distortion (RD) cost calculation, and last nonzero (LNZ) coefficient identification. The proposed optimizations are incorporated into the uvg266 practical VVC encoder. According to our evaluations, these optimization techniques accelerate the DQ process by over 1.7× and the entire VVC intra encoder by 1.4×, without any coding efficiency degradation. The proposed optimizations achieve more than twice the speedup over the existing techniques. Joose Sainio, Alexandre Mercat, Jarno Vanne |
VCIP | 2 |
| 2022 | Spatio-Temporal Parallelization Scheme for HEVC Encoding on Multi-Computer SystemsabstractHigh Efficiency Video Coding (HEVC) sets the scene for economic video transmission and storage, but its inherent computational complexity calls for efficient parallelization techniques. This paper introduces and compares three different parallelization strategies for HEVC encoding on multi-computer systems: 1) spatial parallelization scheme, where input video frames are divided into slices and distributed among available computers; 2) temporal parallelization scheme, where input video is distributed among computers in groups of consecutive frames; 3) spatio-temporal parallelization scheme that combines the proposed spatial and temporal approaches. All these three schemes were benchmarked as part of the practical Kvazaar open-source HEVC encoder. Our experimental results on 2–5 computer configurations show that using the spatial scheme gives 1.65×–2.90× speedup at the cost of 4.16%–13.09% bitrate loss over a single-computer setup. The respective speedup with temporal parallelization is 1.86×–3.26× without any coding overhead. The spatio-temporal scheme with 2 slices was shown to offer the best load-balancing with 1.81×–3.55× speedups and a constant coding loss of 4.16%. Alexandre Mercat, Sami Ahovainio, Jarno Vanne |
ICIP | 1 |
| 2021 | Parallel Implementations of Lambda Domain and R-Lambda Model Rate Control Schemes in a Practical HEVC EncoderabstractThis paper summarizes the key parallelization aspects of the two most popular academic rate control (RC) algorithms for High Efficiency Video Coding (HEVC): 1) Lambda domain (LD) [1] and 2) R-Lambda model (R-LM) [2] that were originally designed for the HEVC test model (HM). In this work, these two sequential RC algorithms were ported to the practical Kvazaar open-source HEVC encoder [3] and adapted to work with wavefront parallel processing (WPP) and overlapping wavefront (OWF) parallelization techniques. Joose Sainio, Alexandre Mercat, Jarno Vanne |
DCC | 2 |
| 2021 | Open-source RTP Library for End-to-End Encrypted Real-Time Video Streaming ApplicationsabstractInformation security has become a key success factor for streaming media applications that are increasingly vulnerable to wiretapping, message forgery, data tampering, hacking, and other possible cyberattacks. This paper addresses the existing security risks in real-time video streaming by introducing a new security extension to our uvgRTP opensource Real-time Transport Protocol (RTP) library. The proposed solution improves content integrity and privacy by adopting Secure RTP (SRTP) and Zimmermann RTP (ZRTP) for media End-to-End Encryption (E2EE). These new security mechanisms make uvgRTP the first open-source library that supports on-the-fly encrypted AVC, HEVC, and VVC video streaming. Our performance results on Intel Core i7-4770 processor show that uvgRTP is able to transport encrypted 8K VVC video at up to 187 fps and 8K HEVC video at up to 120 fps over a 10 Gbps Local Area Network (LAN). The achieved transfer rate for encrypted HEVC video is 50% higher and latency 86% lower than the respective performance values of FFmpeg in unencrypted HEVC streaming. These top streaming speed results with state-of-the-art video codec support, advanced encryption mechanisms, and the permissive BSD license make uvgRTP an attractive solution for a broad range of commercial and academic streaming media applications. Joni Räsänen, Aaro Altonen, Alexandre Mercat, Jarno Vanne |
ISM | 3 |
| 2021 | uvgVenctester: Open-Source Test Automation Framework for Comprehensive Video Encoder BenchmarkingabstractThe agile and efficient development of modern video encoders calls for automated testing methodologies. This paper presents the first-of-its-kind open-source test automation framework called uvgVenctester (github.com/ultravideo/uvgVenctester) that is designed for comprehensive performance and conformance testing of video encoders with the desired set of test video sequences. Our framework comes with built-in support for the popular AVC, HEVC, VVC, VP9, and AV1 video coding formats and the state-of-the-art HM, Kvazaar, x265, VTM, VVenC, SVT-VP9, and SVT-AV1 video encoders. Furthermore, there are no technical limitations of adopting other formats or encoders. The developers can evaluate the encoder of interest under the three primary usage scenarios: 1) conformance testing of the encoded bitstream; 2) rate-distortion-complexity comparison with the other encoders; and 3) systematic exploration of encoding parameters. The framework provides commonly used analysis tools to quantify encoding quality, speed, and bitrate with versatile set of absolute and comparative results such as Bjøntegaard Delta (BD)-Rate for PSNR, SSIM, and VMAF quality metrics. The supported output formats include CSV, graph, and comparison table. They ensure that the results are available in human and machine-readable formats. To the best of our knowledge, the proposed framework is currently the most comprehensive and modular open-source software toolset for video encoder benchmarking. Joose Sainio, Alexandre Mercat, Jarno Vanne |
MMSys | 2 |
| 2020 | Live Demonstration: Multi-Laptop HEVC EncodingabstractThis paper presents a demonstration setup for distributed real-time HEVC encoding on a multi-computer system. The demonstrated multi-level parallelization scheme is implemented in the practical Kvazaar open-source HEVC encoder. It allows Kvazaar to exploit parallelism at three levels: 1) Single Instruction Multiple Data (SIMD) optimized coding tools at the data level; 2) Wavefront Parallel Processing (WPP) and Overlapped Wavefront (OWF) parallelization strategies at the thread level; and 3) distributed slice encoding on multi-computer systems at the process level. This interactive demonstration allows visitors to gradually increase the degree of parallelism in Kvazaar and see the benefits of parallelization in live HEVC encoding. Exploiting all three levels of parallelism on a three-laptop setup speeds up Kvazaar by almost 21× over a non-parallelized single-core implementation of Kvazaar. Sami Ahovainio, Alexandre Mercat, Jarno Vanne |
ISCAS | 2 |
| 2020 | Multi-Level Parallelization Scheme for Distributed HEVC Encoding on Multi-Computer SystemsabstractHigh Efficiency Video Coding (HEVC) creates the conditions for cost-effective video transmission and storage but its inherent computational complexity calls for efficient parallelization techniques. This paper provides HEVC encoders with a holistic parallelization scheme that exploits parallelism at data, thread, and process levels at the same time. The proposed scheme is implemented in the practical Kvazaar open-source HEVC encoder. It makes Kvazaar exploit parallelism at three levels: 1) Single Instruction Multiple Data (SIMD) optimized coding tools at the data level; 2) Wavefront Parallel Processing (WPP) and Overlapped Wavefront (OWF) parallelization strategies at the thread level; and 3) distributed slice encoding on multi-computer systems at the process level. Our results show that the proposed process-level parallelization scheme increases the coding speed of Kvazaar by 1.86× on two computers and up to 3.92× on five computers with +0.19% and +0.81% coding losses, respectively. Exploiting all these three parallelism levels on a five-computer setup gives almost a 25× speedup over a non-parallelized single-core implementation. Sami Ahovainio, Alexandre Mercat, Marko Viitanen, Jarno Vanne |
ISCAS | 2 |
| 2020 | Live Demonstration: Interactive Quality of Experience Evaluation in Kvazzup Video CallabstractThis paper presents an interactive demonstration setup, which allows users to configure the video coding parameters of Kvazzup open-source video call software at runtime and evaluate their impact on Quality of Service (QoS) and Quality of Experience (QoE). The demonstration is carried out by implementing a new Kvazzup control panel for video call parameterization and visual quality, bit rate, latency, and frame rate evaluation. Joni Räsänen, Aaro Altonen, Alexandre Mercat, Jarno Vanne |
ISM | 3 |
| 2020 | Kvazaar 2.0: fast and efficient open-source HEVC inter encoderabstractHigh Efficiency Video Coding (HEVC) is the key to economic video transmission and storage in the current multimedia applications but tackling its inherent computational complexity requires powerful video codec implementations. This paper presents Kvazaar 2.0 HEVC encoder that is the new release of our academic open-source software (github.com/ultravideo/kvazaar). Kvazaar 2.0 introduces novel inter coding functionality that is built on advanced rate-distortion optimization (RDO) scheme and speeded up with several early termination mechanisms, SIMD-optimized coding tools, and parallelization strategies. Our experimental results show that the proposed coding scheme makes Kvazaar 125 times as fast as the HEVC reference software HM on the Intel Xeon E5-2699 v4 22-core processor at the additional coding cost of only 2.4% on average. In constant quantization parameter (QP) coding, Kvazaar is also 3 times as fast as the respective preset of the well-known practical x265 HEVC encoder and is still able to attain 10.7% lower average bit rate than x265 for the same objective visual quality. These results indicate that Kvazaar has become one of the leading open-source HEVC encoders in practical high-efficiency video coding. Ari Lemmetti, Marko Viitanen, Alexandre Mercat, Jarno Vanne |
MMSys | 3 |
| 2020 | UVG dataset: 50/120fps 4K sequences for video codec analysis and developmentabstractThis paper provides an overview of our open Ultra Video Group (UVG) dataset that is composed of 16 versatile 4K (3840×2160) test video sequences. These natural sequences were captured either at 50 or 120 frames per second (fps) and stored online in raw 8-bit and 10-bit 4:2:0 YUV formats. The dataset is published on our website (ultravideo.cs.tut.fi) under a non-commercial Creative Commons BY-NC license. In this paper, all UVG sequences are described in detail and characterized by their spatial and temporal perceptual information, rate-distortion behavior, and coding complexity with the latest HEVC/H.265 and VVC/H.266 reference video codecs. The proposed dataset is the first to provide complementary 4K sequences up to 120 fps and is therefore particularly valuable for cutting-edge multimedia applications. Our evaluations also show that it comprehensively complements the existing 4K test set in VVC standardization, so we recommend including it in subjective and objective quality assessments of next-generation VVC codecs. Alexandre Mercat, Marko Viitanen, Jarno Vanne |
MMSys | 1 |
| 2020 | Tunable VVC Frame Partitioning Based on Lightweight Machine LearningabstractBlock partition structure is a critical module in video coding scheme to achieve significant gap of compression performance. Under the exploration of the future video coding standard, named Versatile Video Coding (VVC), a new Quad Tree Binary Tree (QTBT) block partition structure has been introduced. In addition to the QT block partitioning defined in High Efficiency Video Coding (HEVC) standard, new horizontal and vertical BT partitions are enabled, which drastically increases the encoding time compared to HEVC. In this paper, we propose a lightweight and tunable QTBT partitioning scheme based on a Machine Learning (ML) approach. The proposed solution uses Random Forest classifiers to determine for each coding block the most probable partition modes. To minimize the encoding loss induced by misclassification, risk intervals for classifier decisions are introduced in the proposed solution. By varying the size of risk intervals, tunable trade-off between encoding complexity reduction and coding loss is achieved. The proposed solution implemented in the JEM-7.0 software offers encoding complexity reductions ranging from 30average for only 0.7% to 3.0% Bjxntegaard Delta Rate (BDBR) increase in Random Access (RA) coding configuration, with very slight overhead induced by Random Forest. The proposed solution based on Random Forest classifiers is also efficient to reduce the complexity of the Multi-Type Tree (MTT) partitioning scheme under the VTM-5.0 software, with complexity reductions ranging from 25% to 61% in average for only 0.4% to 2.2% BD-BR increase. Thomas Amestoy, Alexandre Mercat, Wassim Hamidouche, Daniel Ménard, Cyril Bergeron |
IEEE Trans. Image Process. | 2 |
| 2019 | Random Forest Oriented Fast QTBT Frame PartitioningabstractBlock partition structure is a critical module in video coding scheme to achieve significant gap of compression performance. Under the exploration of future video coding standard by the Joint Video Exploration Team (JVET), named Versatile Video Coding (VVC), a new Quad Tree Binary Tree (QTBT) block partition structure has been introduced. In addition to the QT block partitioning defined by High Efficiency Video Coding (HEVC) standard, new horizontal and vertical BT partitions are enabled, which drastically increases the encoding time compared to HEVC. In this paper, we propose a fast QTBT partitioning scheme based on a Machine Learning approach. Complementary to techniques proposed in literature to reduce the complexity of HEVC Quad Tree (QT) partitioning, the propose solution uses Random Forest classifiers to determine for each block which partition modes between QT and BT is more likely to be selected. Using uncertainty zones of classifier decisions, the proposed complexity reduction technique is able to reduce in average by 30% the encoding time of JEM-v7.0 software in Random Access configuration with only 0.57% Bjøntegaard Delta Rate (BD-BR) increase. Thomas Amestoy, Alexandre Mercat, Wassim Hamidouche, Cyril Bergeron, Daniel Ménard |
ICASSP | 2 |
| 2019 | Convex Energy Optimization of Streaming Applications for MPSoCsabstractThe energy efficiency of modern MPSoCs is enhanced by complex hardware features such as Dynamic Voltage and Frequency Scaling (DVFS) and Dynamic Power Management (DPM). This paper introduces a new method, based on convex problem solving, that determines the most energy efficient operating point in terms of frequency and number of active cores in an MPSoC. The solution can challenge the popular approaches based on never-idle (or As-Slow-As-Possible (ASAP)) and race-to-idle (or As-Fast-As-Possible (AFAP)) principles. Experimental data are reported using a Samsung Exynos 5410 MPSoC and show a reduction in energy of up to 27 % when compared to ASAP and AFAP. Erwan Nogues, Alexandre Mercat, Florian Arrestier, Maxime Pelcat, Daniel Ménard |
ICASSP | 2 |
| 2019 | Acceleration of Kvazaar HEVC Intra Encoder With Machine LearningabstractThe complexity of High Efficiency Video Coding (HEVC) poses a real challenge to HEVC encoder implementations. Particularly, the complexity stems from the HEVC quad-tree structure that also has an integral part in HEVC coding efficiency. This paper presents a Machine Learning (ML) based technique for pruning the HEVC quad-tree without deteriorating coding gain. We show how ML decision trees can be used to predict a depth interval for a quad-tree before the Rate-Distortion Optimization (RDO). This approach limits the number of RDO candidates and thus speeds up encoding. The proposed technique works particularly well with high-quality video coding and it is shown to accelerate the veryslow preset of practical Kvazaar HEVC intra encoder by 1.35× with 0.49% bit rate increase. Compared with the corresponding preset of x265 encoder, Kvazaar is 2.12× as fast at a cost of under 1.21% bit rate overhead. These results indicate that the optimized Kvazaar is the leading open-source encoder in high-quality HEVC intra coding. Alexandre Mercat, Ari Lemmetti, Marko Viitanen, Jarno Vanne |
ICIP | 1 |
| 2019 | Remote VR Gaming on Mobile DevicesabstractThis paper presents a remote 360-degree virtual reality (VR) gaming system for mobile devices. In this end-to-end scheme, execution of VR game is off-loaded from low-power mobile devices to a remote server where the executed game is rendered based on controller orientation and actions transmitted over the network. The server is running the Unity game engine and Kvazaar video encoder. Kvazaar compresses the rendered views of the game to High Efficiency Video Coding (HEVC) video that is streamed to a player over a regular WiFi link in real time. The frontend of our proof-of-concept demonstrator setup is composed of the Samsung Galaxy S8 smartphone and Google Daydream View VR headset with a controller. The backend server is a laptop equipped with Nvidia GTX 1070 GPU and Intel i7 7820HK CPU. The system is able to run the demonstrated 360-degree shooting VR game with 1080p resolution at 30 fps while keeping motion-to-photon latency close to 50 ms. This approach lets players enter immersive gaming experience without a need to invest in all-in-one VR headsets. Mikko Juhani Pitkänen, Marko Viitanen, Alexandre Mercat, Jarno Vanne |
ACM Multimedia | 3 |
| 2019 | Complexity Reduction Opportunities in the Future VVC Intra EncoderabstractThe Joint Video Expert Team (JVET) is developing the next-generation video coding standard called Versatile Video Coding (VVC) and their ultimate goal is to double the coding efficiency over the current state-of-the-art standard HEVC without letting complexity get out of hand. This work addresses the complexity of the VVC reference encoder called VVC Test Model (VTM) under All Intra coding configuration. The VTM3.0 is able to improve intra coding efficiency by 21% over the latest HEVC reference encoder HM16.19. This coding gain primarily stems from three new coding tools. First, the HEVC Quad-Tree (QT) structure extension with Multi-Type Tree (MTT) partitioning. Second, the duplication of intra prediction modes from 35 to 67. And third, the Multiple Transform Selection (MTS) scheme with two new discrete cosine/sine transforms (DCT-VIII and DST-VII). However, these new tools also play an integral part in making VTM intra encoding around 20 times as complex as that of HM. The purpose of this work is to analyze these tools individually and specify theoretical upper limits for their complexity reduction. According to our evaluations, the complexity reduction opportunity of block partitioning is up to 97%, i.e., the encoding complexity would drop down to 3% for the same coding efficiency if the optimal block partitioning could be directly predicted. The respective percentages for intra mode reduction and MTS optimization are 65% and 55%. We believe these results motivate VVC codec designers to develop techniques that are able to take most out of these opportunities. Alexandre Tissier, Alexandre Mercat, Thomas Amestoy, Wassim Hamidouche, Jarno Vanne, Daniel Ménard |
MMSP | 2 |
| 2019 | Public and open HEVC encoding service in the cloudabstractThe ability to record vast amounts of video content requires convenient and efficient video coding services with which users can tackle the limited storage and transmission capacities. This paper presents an open-source cloud service for encoding raw video formats and transcoding compressed videos to the latest HEVC/H.265 format. Respective commercial transcoding services are available on the Internet but they are behind a paywall. On the other hand, using command-line interfaces of existing open-source software solutions requires in-depth knowledge of the coding process to attain the best coding gain and speed. The proposed service is available online, it is free to use without any registration, and its easy-to-use web interface makes it feasible for non-technical users. It is built on the FFmpeg multimedia framework whose built-in decoders accept various input video formats that are then compressed to HEVC with a full-fledged Kvazaar open-source encoder. Aaro Altonen, Marko Viitanen, Joni Räsänen, Alexandre Mercat, Jarno Vanne |
MMSys | 4 |
| 2019 | Parallax-Tolerant 360 Live Video StitcherabstractThis paper presents an open-source software implementation for real-time 360-degree video stitching. To ensure a seamless stitching result, cylindrical and content-preserving warping are implemented to dynamically correct image alignment and parallax, which may drift due to scene changes, moving objects, or camera movement. Depth variation, color changes, and lighting differences between adjacent frames are also smoothed out to improve visual quality of the panoramic video. The system is benchmarked with six 1080p videos, which are stitched into 4096×732 pixel output format. The proposed algorithm attains an output rate of 18 frames per second on GeForce GTX 1070 GPU and real-time speed can be met with a high-end GPU. Miko Atokari, Marko Viitanen, Alexandre Mercat, Emil Kattainen, Jarno Vanne |
VCIP | 3 |
| 2018 | Machine Learning Based Choice of Characteristics for the One-Shot Determination of the HEVC Intra Coding TreeabstractIn the last few years, the Internet of Things (IoT) has become a reality. Forthcoming applications are likely to boost mobile video demand to an unprecedented level. A large number of systems are likely to integrate the latest MPEG video standard High Efficiency Video Coding (HEVC) in the long run and will particularly require energy efficiency. In this context, constraining the computational complexity of embedded HEVC encoders is a challenging task, especially in the case of software encoders. The most energy consuming part of a software intra encoder is the determination of the coding tree partitioning, i.e. the size of pixel blocks. This determination usually requires an iterative process that leads to repeating some encoding tasks. State-of-the-art studies have focused on predicting, from “easily” computed characteristics, an efficient coding tree. They have proposed and evaluated independently many characteristics for one-shot quad-tree prediction. In this paper, we present a fair comparison of these characteristics using a Machine Learning approach and a real-time HEVC encoder. Both computational complexity and information gain are considered, showing that characteristics are far from equivalent in terms of coding tree prediction performance. Alexandre Mercat, Florian Arrestier, Maxime Pelcat, Wassim Hamidouche, Daniel Ménard |
PCS | 1 |
| 2018 | Reproducible Evaluation of System Efficiency With a Model of Architecture: From Theory to PracticeabstractCurrent trends in high performance and embedded computing include design of increasingly complex hardware architectures with high parallelism, heterogeneous processing elements, and nonuniform communication resources. In order to take hardware and software design decisions, early evaluations of the system nonfunctional properties are needed. These evaluations of system efficiency require electronic system-level information on both algorithms and architecture. Contrary to algorithm models for which a major body of work has been conducted on defining formal models of computation (MoCs), architecture models from the literature are mostly empirical models from which reproducible experimentation requires the accompanying software. In this paper, a precise definition of a model of architecture (MoA) is proposed that focuses on reproducibility and abstraction and removes the overlap previously existing between the notions of MoA and MoC. A first MoA, called the linear system-level architecture model (LSLA), is presented. To demonstrate the generic nature of the proposed new architecture modeling concepts, we show that the LSLA model can be integrated flexibly with different MoCs. LSLA is then used to model the energy consumption of a state-of-the-art multiprocessor system-on-chip (MPSoC) when running an application described using the synchronous dataflow MoC. A method to automatically learn LSLA model parameters from platform measurements is introduced. Despite the high complexity of the underlying hardware and software, a simple LSLA model is demonstrated to estimate the energy consumption of the MPSoC with a fidelity of 86%. Maxime Pelcat, Alexandre Mercat, Karol Desnos, Luca Maggiani, Yanzhou Liu 0001, Julien Heulot, Jean-François Nezan, Wassim Hamidouche, Daniel Ménard, Shuvra S. Bhattacharyya |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | Exploiting computation skip to reduce energy consumption by approximate computing, an HEVC encoder case studyabstractApproximate computing paradigm provides methods to optimize algorithms with considering both computational accuracy and complexity. This paradigm can be exploited at different levels of abstraction, from technological to application levels. Approximate computing at algorithm level aims at reducing computational complexity by approximating or skipping block functions of the computation. Numerous applications in the signal and image processing domain integrate algorithms based on discrete optimization techniques. These techniques minimize a cost function by exploring the search space. In this paper, a new approach is proposed to exploit the computation-skipping approximate computing concept by using the Smart Search Space Reduction (Sssr) technique. Sssr enables early selection of the best candidate configurations to reduce the search space. An efficient SSSR technique adjusts configuration selectivity to reduce execution complexity while selecting the most suitable functions to skip. The High Efficiency Video Coding (HEVC) encoder in All Intra (AI) profile is used as a case study to illustrate the benefits of SSSR. In this application, two functions use discrete optimization to explore different solutions and select the one leading to the minimal cost in terms of bitrate/quality and computational energy: coding-tree partitioning and intra-mode prediction. By applying SSSR to this use case, energy reductions from 20% to 70% are explored through Pareto in Rate-Energy space. Alexandre Mercat, Justine Bonnot, Maxime Pelcat, Wassim Hamidouche, Daniel Ménard |
DATE | 1 |
| 2017 | Energy reduction opportunities in an HEVC real-time encoderabstractHigh Efficiency Video Coding (HEVC) is one of the latest released video standards and offers up to 40% bitrate savings when compared to the widespread H.264/AVC standard, at the cost of a substantial complexity growth. Constraining the complexity of HEVC encoding is a challenging task for embedded applications based on a software encoder. In the last few years, the Internet of Thingss (IoTs) has become a reality. Forecoming applications are likely to boost mobile video demand to an unprecedented level. In this context, designing energy-efficient HEVC real-time encoders is becoming a major challenge for software and hardware designers. In this paper, an analysis is conducted of the energy reduction opportunities offered by an HEVC encoder. The energy reduction search space is demonstrated, and the impact on energy consumption of encoding tools at various levels of granularity is measured. Alexandre Mercat, Florian Arrestier, Wassim Hamidouche, Maxime Pelcat, Daniel Ménard |
ICASSP | 1 |
| 2017 | Constrain the Docile CTUs: An In-Frame complexity allocator for HEVC Intra encodersabstractHigh Efficiency Video Coding (HEVC) is one of the latest released video standards and offers up to 40% bitrate savings when compared to the widespread H.264/AVC standard, at the cost of a substantial complexity growth. Constraining the complexity of HEVC encoding is a challenging task for embedded applications based on a software encoder. The most frequent approach to solve this problem is to optimise the coding tree structure to balance compression efficiency and computational complexity. In this context, we propose and assess a method to adequately allocate the computational complexity among coding units in a frame encoded in Intra mode. By studying an open-source real-time HEVC encoder, correlations are observed between Rate-Distortion (RD)-cost and encoding complexity that motivate a new complexity allocation technique. This technique, called “Constrain the Docile CTUs” (CDC), consists of allocating less computational complexity to units with low RD-costs and using RD-costs from preceding images as predictors for the current RD-costs. Experimental results demonstrate substantial gains, up to 36% of Bjøntegaard Delta Bit Rate (BD-BR), when using CDC method instead of other allocation methods. Alexandre Mercat, Florian Arrestier, Wassim Hamidouche, Maxime Pelcat, Daniel Ménard |
ICASSP | 1 |
| 2017 | Smart search space reduction for approximate computing: A low energy HEVC encoder case study
Alexandre Mercat, Justine Bonnot, Maxime Pelcat, Karol Desnos, Wassim Hamidouche, Daniel Ménard |
J. Syst. Archit. | 1 |
| 2016 | Optimized Belief Propagation Algorithm onto Embedded Multi and Many-Core Systems for Stereo MatchingabstractStereo matching techniques aim at reconstructing disparity maps from a pair of images. The use of stereo matching techniques in embedded systems is very challenging due to the complexity of the state-of-the-art algorithms. Local stereo matching algorithms are efficiently implemented on GPU and DSP. This paper presents the optimization of the One Dimension Belief Propagation (BP-1D) algorithm. BP-1D is faster than previous algorithms on monocore DSP and its implementation onto multicore DSPs is straightforward. BP-1D implemented on multicore embedded platforms out-performs previous stereo matching implementations reaching real-time performances for resolutions up to 1080p with a 10 Watts power consumption. Jean-François Nezan, Alexandre Mercat, Patrice Delmas, Georgy L. Gimel'farb |
PDP | 2 |
| 2016 | Energy Efficient Scheduling of Real Time Signal Processing Applications through Combined DVFS and DPMabstractThis paper proposes a framework to design energy efficient signal processing systems. The energy efficiency is provided by combining Dynamic Frequency and Voltage Scaling (DVFS) and Dynamic Power Management (DPM). The framework is based on Synchronous Dataflow (SDF) modeling of signal processing applications. A transformation to a single rate form is performed to expose the application parallelism. An automated scheduling is then performed, minimizing the constraint of energy efficiency and providing DVFS and DPM decisions. This framework uses an architecture model including the number of available cores, the per-actor processing load and the energy per-cycle, derived from time and power measurements of modelled applications. After introducing the proposed framework, the energy characterization of big.LITTLE SoC systems is described. A generic approach is presented to generate the energy model of a platform from power measurements as customized polynomials. Finally, the experimental results on a Samsung Exynos 5410 big.LITTLE processor show that the energy optimal execution is not obtained by Linux governors that can execute either as-fast-as-possible or as-slow-as-possible. Instead, the most energy efficient scheduling is obtained by adapting both DVFS and DPM to application needs. Erwan Nogues, Maxime Pelcat, Daniel Ménard, Alexandre Mercat |
PDP | 4 |
| 2014 | Implementation of a Stereo Matching algorithm onto a Manycore Embedded SystemabstractStereo Matching techniques aim at reconstructing the disparity maps with a pair of images. The use of Stereo Matching techniques in embedded systems is very challenging due to the complexity of the state of the art algorithms. This paper proposes a real-time Stereo Matching algorithm optimised for the last generation of Manycore Embedded Systems. The features and parameters of the algorithms have been chosen to optimise the trade-off between an high quality and a low complexity. A memory analysis is performed for the algorithm's decomposition and the resulting mapping on the Manycore platform is provided. Algorithm and arithmetic optimisations have been applied to each part of the algorithm to decrease the execution time down to 160ms for a CIF resolution. Alexandre Mercat, Jean-François Nezan, Daniel Ménard |
ISCAS | 1 |