EDBT 2026 Demo / reviewers in the wild / expert
Ari Lemmetti
dblp:166/3426
· DBLP profile ↗
11ranked-venue papers
3as first author
3since 2021 · last 2023
0000-0002-2798-4893ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-authorSystems, architecture and hardware · 4 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | AVX2-Optimized Interpolation Filters for HEVC Inter EncodingabstractHigh Efficiency Video Coding (HEVC) sets the stage for economic video transmission and storage, but its inherent computational complexity calls for powerful implementations. This paper addresses the principal performance bottleneck of HEVC codecs by introducing AVX2-vectorized algorithms for HEVC interpolation filters. The proposed speed-up techniques include 1) a data permutation scheme for the horizontal interpolation stage; 2) a sliding window strategy for the vertical interpolation stage; 3) optimal usage of horizontal and vertical interpolation during fractional motion estimation; and 4) a lane-based approach to double the vector lengths from 128-bit legacy vector extensions to 256bits of AVX2. Our AVX2-optimized interpolation filters were benchmarked as part of the practical Kvazaar open-source HEVC encoder. On an Intel 8-core Xeon processor, they were shown to be 9.7 and 8.5 times as fast as scalar interpolation with the Kvazaar ultrafast and veryslow presets, respectively. In both cases, changing over from scalar to vectorized interpolation increases the coding speed of Kvazaar by more than 50%, which stresses the importance of interpolation optimizations in modern video encoders. Alexandre Mercat, Ari Lemmetti, Joose Sainio, Jarno Vanne |
ISCAS | 2 |
| 2022 | High-Level Synthesis Implementation of an Embedded Real-Time HEVC Intra Encoder on FPGA for Media ApplicationsabstractHigh Efficiency Video Coding (HEVC) is the key enabling technology for numerous modern media applications. Overcoming its computational complexity and customizing its rich features for real-time HEVC encoder implementations, calls for automated design methodologies. This article introduces the first complete High-Level Synthesis (HLS) implementation for HEVC intra encoder on FPGA. The C source code of our open-source Kvazaar HEVC encoder is used as a design entry point for HLS that is applied throughout the whole encoder design process, from data-intensive coding tools like intra prediction and discrete transforms to more control-oriented tools such as context-adaptive binary arithmetic coding (CABAC). Our prototype is run on Nokia AirFrame Cloud Server equipped with 2.4 GHz dual 14-core Intel Xeon processors and two Intel Arria 10 PCIe FPGA accelerator cards with 40 Gigabit Ethernet. This proof-of-concept system is designed for hardware-accelerated HEVC encoding and it achieves real-time 4K coding speed up to 120 fps. The coding performance can be easily scaled up by adding practically any number of network-connected FPGA cards to the system. These results indicate that our HLS proposal not only boosts development time, but also provides previously unseen design scalability with competitive performance over the existing FPGA and ASIC encoder implementations. Panu Sjövall, Ari Lemmetti, Jarno Vanne, Sakari Lahti, Timo Hämäläinen 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2021 | High-Level Synthesis Implementation of Transform-Exempted SATD Architectures for Low-Power Video CodingabstractThis paper presents the first known high-level synthesis (HLS) implementation for the Sum of Absolute Transformed Differences (SATD) calculation. The proposed hardware architecture is designed for two SATD algorithms: a widespread Fast Walsh-Hadamard Transform (FWHT-SATD) and a recently introduced Transform Exempted scheme (TE- SATD). This 2-stage architecture is made up of two 1-D Walsh- Hadamard Transform (WHT) stages and a transpose buffer (TB) between them. The chosen HLS approach cuts down design time over contemporary design methods and thereby made it feasible to implement a set of dedicated FWHT-SATD and TE- SATD architectures for 4x4, 8x8, and 16x16 pixel blocks. All these six architectures were synthesized for 28 nm and 45 nm standard cell technologies, and their area and energy consumptions were analysed. TE-based implementations provide 6.0-8.3% total cell area savings and 6.9-12.7% better energy-efficiency than traditional FWHT approaches. Our proposal is the first to introduce TE-SATD architectures for up to 16x16 blocks and each of these tailored architectures was shown to provide better trade-off between silicon area and performance than their reference implementations. Tero Partanen, Ari Lemmetti, Panu Sjövall, Jarno Vanne |
ISCAS | 2 |
| 2020 | Real-Time Implementation Of Scalable Hevc EncoderabstractThis paper presents the first known open-source Scalable HEVC (SHVC) encoder for real-time applications. Our proposal is built on top of Kvazaar HEVC encoder by extending its functionality with spatial and signal-to-noise ratio (SNR) scalable coding schemes. These two scalability schemes have been optimized for real-time coding by means of three parallelization techniques: 1) wavefront parallel processing (WPP); 2) overlapped wavefront (OWF); and 3) AVX2-optimized upsampling. On an 8-core Xeon W-2145 processor, the proposed spatially scalable Kvazaar can encode twolayer 1080p video above 50 fps with scaling ratios of l.5 and 2. The respective coding gain s are 18.4% and 9.9% over Kvazaar simulcast coding at similar speed. Correspondingly, the coding speed of SNR scalable Kvazaar exceeds 30 fps with two-layer 1080p video. On average, it obtain s1.20 times speedup and 17.0% better coding efficiency over the simulcast case. These results justify the benefits of the proposed scalability schemes in real-time SHVC coding. Jaakko Laitinen, Ari Lemmetti, Jarno Vanne |
ICIP | 2 |
| 2020 | Kvazaar 2.0: fast and efficient open-source HEVC inter encoderabstractHigh Efficiency Video Coding (HEVC) is the key to economic video transmission and storage in the current multimedia applications but tackling its inherent computational complexity requires powerful video codec implementations. This paper presents Kvazaar 2.0 HEVC encoder that is the new release of our academic open-source software (github.com/ultravideo/kvazaar). Kvazaar 2.0 introduces novel inter coding functionality that is built on advanced rate-distortion optimization (RDO) scheme and speeded up with several early termination mechanisms, SIMD-optimized coding tools, and parallelization strategies. Our experimental results show that the proposed coding scheme makes Kvazaar 125 times as fast as the HEVC reference software HM on the Intel Xeon E5-2699 v4 22-core processor at the additional coding cost of only 2.4% on average. In constant quantization parameter (QP) coding, Kvazaar is also 3 times as fast as the respective preset of the well-known practical x265 HEVC encoder and is still able to attain 10.7% lower average bit rate than x265 for the same objective visual quality. These results indicate that Kvazaar has become one of the leading open-source HEVC encoders in practical high-efficiency video coding. Ari Lemmetti, Marko Viitanen, Alexandre Mercat, Jarno Vanne |
MMSys | 1 |
| 2019 | Acceleration of Kvazaar HEVC Intra Encoder With Machine LearningabstractThe complexity of High Efficiency Video Coding (HEVC) poses a real challenge to HEVC encoder implementations. Particularly, the complexity stems from the HEVC quad-tree structure that also has an integral part in HEVC coding efficiency. This paper presents a Machine Learning (ML) based technique for pruning the HEVC quad-tree without deteriorating coding gain. We show how ML decision trees can be used to predict a depth interval for a quad-tree before the Rate-Distortion Optimization (RDO). This approach limits the number of RDO candidates and thus speeds up encoding. The proposed technique works particularly well with high-quality video coding and it is shown to accelerate the veryslow preset of practical Kvazaar HEVC intra encoder by 1.35× with 0.49% bit rate increase. Compared with the corresponding preset of x265 encoder, Kvazaar is 2.12× as fast at a cost of under 1.21% bit rate overhead. These results indicate that the optimized Kvazaar is the leading open-source encoder in high-quality HEVC intra coding. Alexandre Mercat, Ari Lemmetti, Marko Viitanen, Jarno Vanne |
ICIP | 2 |
| 2018 | Rate-Distortion-Complexity Optimized Coding Scheme for Kvazaar HEVC Intra EncoderabstractThis paper summarizes a low-complexity rate-distortion optimization (RDO) scheme for Kvazaar HEVC intra encoder (github.com/ultravideo/kvazaar). Our work particularly addresses RDO quantization (RDOQ) since it is the most complex intra coding tool taking almost 60% of the Kvazaar complexity. Ari Lemmetti, Eemeli Kallio, Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001 |
DCC | 1 |
| 2017 | Kvazaar: HEVC/H.265 4K30p Intra EncoderabstractThis paper demonstrates the usage of Kvazaar open-source HEVC intra encoder in 4K real-time video encoding. In this setup, a raw 4K video is shot by an action camera, captured by an HDMI capture card, encoded in real-time by Kvazaar ultrafast preset on a 22-core Intel Xeon processor, sent to a laptop, and decoded by OpenHEVC decoder for playback. The encoding process is visualized on the fly by Kvazaar run-time visualizer. Arttu Ylä-Outinen, Ari Lemmetti, Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001 |
ISM | 2 |
| 2016 | AVX2-optimized Kvazaar HEVC intra encoderabstractThis paper presents efficient SIMD optimizations for the open-source Kvazaar HEVC intra encoder. The C implementation of Kvazaar is accelerated by Intel AVX2 instructions whose effect on Kvazaar ultrafast preset is profiled. According to our profiling results, C functions of SATD, DCT, quantization, and intra prediction account for over 60% of the total intra coding time of Kvazaar ultrafast preset. This work shows that optimizing primarily these functions doubles the coding speed of a single-threaded Kvazaar intra encoder for the same rate-distortion performance. The highest performance boost is obtained by deploying the proposed optimizations jointly with multithreading. On the Intel 8-core i7 processor, the AVX2-optimized 16-threaded Kvazaar ultrafast preset achieves real-time (30 fps) intra coding speed up to 1080p resolution. Compared to AVX2-optimized ultrafast preset of x265, Kvazaar is 20% times faster and still obtains 9.1% bit rate gain for the same quality. These results justify that Kvazaar is currently the leading open-source HEVC intra encoder in terms of real-time coding speed and efficiency. Ari Lemmetti, Ari Koivula, Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001 |
ICIP | 1 |
| 2016 | Kvazaar: Open-Source HEVC/H.265 EncoderabstractKvazaar is an academic software video encoder for the emerging High Efficiency Video Coding (HEVC/H.265) standard. It provides students, academic professionals, and industry experts a free, cross-platform HEVC encoder for x86, x64, PowerPC, and ARM processors on Windows, Linux, and Mac. Kvazaar is being developed from scratch in C and optimized in Assembly under the LGPLv2.1 license. The development is being coordinated by Ultra Video Group at Tampere University of Technology (TUT) and the implementation work is carried out by an active community on GitHub. Developer friendly source code of Kvazaar makes joining easy for new developers. Currently, Kvazaar includes all essential coding tools of HEVC and its modular source code facilitates parallelization on multi and manycore processors as well as algorithm acceleration on hardware. Kvazaar is able to attain real-time HEVC coding speed up to 4K video on an Intel 14-core Xeon processor. Kvazaar is also supported by FFmpeg and Libav. These de-facto standard multimedia frameworks boost Kvazaar popularity and enable its joint usage with other well-known multimedia processing tools. Nowadays, Kvazaar is an integral part of teaching at TUT and it has got a key role in three Eureka Celtic-Plus projects in the fields of 4K TV broadcasting, virtual advertising, Video on Demand, and video surveillance. Marko Viitanen, Ari Koivula, Ari Lemmetti, Arttu Ylä-Outinen, Jarno Vanne, Timo Hämäläinen 0001 |
ACM Multimedia | 3 |
| 2015 | Kvazaar HEVC encoder for efficient intra codingabstractThis paper presents an open-source Kvazaar encoder for HEVC intra coding. This academic software encoder has been developed from the scratch using C as an implementation language by prioritizing modularity, portability, and readability of the source code. Kvazaar implements almost the same intra coding functionality as HEVC reference encoder (HM) but its rewritten source code makes it significantly faster. In all-intra (AI) coding, a single-threaded C implementation of Kvazaar is 2.3 times faster than HM at a cost of 1.7% bit rate increase. The respective values with a high speed preset of Kvazaar are 10.6 and 8.8%. Compared to a single-threaded C++ implementation of x265, Kvazaar improves rate-distortion performance and increases encoding speed in both high-quality and high-speed test cases. Kvazaar has a particular edge in the high-speed test case where it almost halves the BD-rate loss and more than doubles the performance. Marko Viitanen, Ari Koivula, Ari Lemmetti, Jarno Vanne, Timo Hämäläinen 0001 |
ISCAS | 3 |