EDBT 2026 Demo / reviewers in the wild / expert
Ramin Ghaznavi Youvalari
dblp:193/6795
· DBLP profile ↗
20ranked-venue papers
9as first author
12since 2021 · last 2025
0000-0001-7260-0599ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 9 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Convolutional Cross-Component Models for Chroma Prediction in Video CodingabstractIn this paper we present two novel approaches for improving intra and inter chroma prediction in video coding. Our research demonstrates that treating the cross-component predictor as a two-dimensional convolutional model can significantly enhance chroma prediction performance. The proposed two convolutional models incorporate multiple spatial neighbors, a bias term, and a nonlinear term. For intra-coded blocks, we derive the model coefficients on the reconstructed neighborhood of the block, while for inter-coded blocks, the model coefficients are determined using prediction samples. To evaluate our methods, we implemented them on top of the ECM software that is currently under exploration by the ITU-T/ISO/IEC Joint Video Experts Team. Our intra cross-component predictor achieves BD-rate savings of {−1.47%, −2.90%, −3.02%}, {−0.92%, −2.04%, −2.32%} (Y, U, V) for the all intra and the random access configurations over ECM-5.0, respectively. Our inter cross-component predictor achieves BD-rate savings of {−0.09%, −1.25%, −1.46%}, {−0.04%, −3.42%, −3.85%} for the random access and the low-delay B configurations over ECM-9.0, respectively. Both proposed methods have been adopted into the ECM software. Pekka Astola, Alireza Aminlou, Ramin Ghaznavi Youvalari, Jani Lainema |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Overfitting NN loop-filters in video codingabstractOverfitting is usually regarded as a negative condition since it impairs the generalisation power of a model. Nevertheless, overfitting a Neural Network (NN) on test data may be advantageous to improve the compression efficiency of image/video coding tools and systems. Previous research has demonstrated the benefits of NN overfitting for post-processing operations, i.e. post-filters, but not yet for actual decoding tools. Generally, the NN is overfitted on test data at the encoder end, and the weight update is coded and sent to the decoder end along the image/video bitstream. The proposed approach follows this strategy. In particular, the overfitting of the Low Operation Point (LOP) loop-filter in NN-based Video Coding (NNVC) software is studied. The overall approach yields Bjøntegaard Delta rate (BD-rate) of -7.74%, -13.73% and -12.49%, for the Y, U and V components, respectively. Out of these coding gains, 1.21%, 6.43% and 5.52%, for the Y, U and V components, are attributed to the overfitting. The boost in the coding gains comes with only 1.5% more complexity, due to the multiplier parameters introduced during the overfitting. Ruiying Yang, María Santamaría 0001, Francesco Cricri, Honglei Zhang 0001, Jani Lainema, Ramin Ghaznavi Youvalari, Miska M. Hannuksela, Tapio Elomaa |
VCIP | 6 |
| 2022 | Content-Adaptive Neural Network Post-Processing Filter with NNR-Coded Weight-UpdatesabstractNeural Network (NN) filters improve the perceptual quality of reconstructed videos by reducing compression artefacts. For content adaptation, a few NN-filters use over-fitting. As the adaptation signal is a weight-update, compression is required to minimise significant bitrate overheads. Most approaches, however, use generic data compression algorithms, which are inadequate for coding NN weight-updates. This work introduces a content-adaptive NN post-processing filter with weight-updates coded using the Neural Network compression and Representation (NNR) standard. The bitrate overhead is further decreased by over-fitting only a subset of weights, selected via energy-based analysis. The proposed filter saved about 4.57% (Y), 10.33% (Cb), 6.53% (Cr) Bjøntegaard Delta rate (BD-rate) on top of the Versatile Video Coding (VVC) Test Model (VTM) 11.0 with NN-based Video Coding (NNVC) 1.0, in Random Access (RA) configuration. Compared to the non-over-fitted NN, the performance was doubled; and compared to 7z, NNR reduced the bitrate of the weight-update by ∼64%. María Santamaría 0001, Francesco Cricri, Jani Lainema, Ramin Ghaznavi Youvalari, Honglei Zhang 0001, Miska M. Hannuksela |
ICIP | 4 |
| 2022 | Bridging the Gap Between Image Coding for Machines and HumansabstractImage coding for machines (ICM) aims at reducing the bitrate required to represent an image while minimizing the drop in machine vision analysis accuracy. In many use cases, such as surveillance, it is also important that the visual quality is not drastically deteriorated by the compression process. Recent works on using neural network (NN) based ICM codecs have shown significant coding gains against traditional methods; however, the decompressed images, especially at low bitrates, often contain checkerboard artifacts. We propose an effective decoder finetuning scheme based on adversarial training to significantly enhance the visual quality of ICM codecs, while preserving the machine analysis accuracy, without adding extra bitcost or parameters at the inference phase. The results show complete removal of the checkerboard artifacts at the negligible cost of −1.6% relative change in task performance score. In the cases where some amount of artifacts is tolerable, such as when machine consumption is the primary target, this technique can enhance both pixel-fidelity and feature-fidelity scores without losing task performance. Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Hamed Rezazadegan Tavakoli, Emre Aksu, Miska M. Hannuksela, Esa Rahtu |
ICIP | 4 |
| 2022 | Low-precision post-filtering in video codingabstractNeural Networks (NNs) have demonstrated their effectiveness in tackling challenges involving multimedia content. In the video coding field, NNs are actively exploited as novel tools that complement conventional signal processing tools, as well as end-to-end coding solutions. Since NNs use commonly floating-point arithmetic, different results may be generated in different computing environments, leading to discrepancies and even corrupted reconstructions. Accordingly, this issue is solved by employing fixed-point arithmetic instead. This paper studies the quantisation of a 32-bit floating-point (float32) NN post-filter to 32-bit fixed-point (int32) and 16-bit fixed-point (int16). On top of the VVC Test Model (VTM) 11.0 NN-based Video Coding (NNVC) 1.0, the coding gains of the float32 post-filter are 5.01% (Y), 18.95% (Cb) and 17.33% Cr. Compared to the float32 inference, the quantised models produce coding losses: 0.01% (Y, Cb and Cr) for the int32 inference and 0.45% (Y), 1.97% (Cb) and 1.19% (Cr) for the int16 inference. Nevertheless, the fixed-point approaches achieve bit exact matches in different computing environments. Moreover, the decoding time with int16 is about half the decoding time of int32. Ruiying Yang, María Santamaría 0001, Francesco Cricri, Honglei Zhang 0001, Jani Lainema, Ramin Ghaznavi Youvalari, Miska M. Hannuksela |
ISM | 6 |
| 2021 | Image Coding For Machines: an End-To-End Learned ApproachabstractOver recent years, deep learning-based computer vision systems have been applied to images at an ever-increasing pace, oftentimes representing the only type of consumption for those images. Given the dramatic explosion in the number of images generated per day, a question arises: how much better would an image codec targeting machine-consumption perform against state-of-the-art codecs targeting human-consumption? In this paper, we propose an image codec for machines which is neural network (NN) based and end-to-end learned. In particular, we propose a set of training strategies that address the delicate problem of balancing competing loss functions, such as computer vision task losses, image distortion losses, and rate loss. Our experimental results show that our NN-based codec outperforms the state-of-the-art Versa-tile Video Coding (VVC) standard on the object detection and instance segmentation tasks, achieving -37.87% and -32.90% of BD-rate gain, respectively, while being fast thanks to its compact size. To the best of our knowledge, this is the first end-to-end learned machine-targeted image codec. Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Esa Rahtu |
ICASSP | 4 |
| 2021 | Learned Image Coding for Machines: A Content-Adaptive ApproachabstractToday, according to the Cisco Annual Internet Report (2018-2023), the fastest-growing category of Internet traffic is machine-to-machine communication. In particular, machine-to-machine communication of images and videos represents a new challenge and opens up new perspectives in the context of data compression. One possible solution approach consists of adapting current human-targeted image and video coding standards to the use case of machine consumption. Another approach consists of developing completely new compression paradigms and architectures for machine-to-machine communications. In this paper, we focus on image compression and present an inference-time content-adaptive fine-tuning scheme that optimizes the latent representation of an end-to-end learned image codec, aimed at improving the compression efficiency for machine-consumption. The conducted experiments targeting instance segmentation task network show that our online finetuning brings an average bitrate saving (BD-rate) of -3.66% with respect to our pretrained image codec. In particular, at low bitrate points, our proposed method results in a significant bitrate saving of -9.85%. Overall, our pretrained-and-then-finetuned system achieves - 30.54% BD-rate over the state-of-the-art image/video codec Versatile Video Coding (VVC) on instance segmentation. Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Hamed Rezazadegan Tavakoli, Esa Rahtu |
ICME | 4 |
| 2021 | Learned Enhancement Filters for Image Coding for MachinesabstractMachine-To-Machine (M2M) communication applications and use cases, such as object detection and instance segmentation, are becoming mainstream nowadays. As a consequence, majority of multimedia content is likely to be consumed by machines in the coming years. This opens up new challenges on efficient compression of this type of data. Two main directions are being explored in the literature, one being based on existing traditional codecs, such as the Versatile Video Coding (VVC) standard, that are optimized for human-targeted use cases, and another based on end-to-end trained neural networks. However, traditional codecs have significant benefits in terms of interoperability, real-time decoding, and availability of hardware implementations over end-to-end learned codecs. Therefore, in this paper, we propose learned post-processing filters that are targeted for enhancing the performance of machine vision tasks for images reconstructed by the VVC codec. The proposed enhancement filters provide significant improvements on the target tasks compared to VVC coded images. The conducted experiments show that the proposed post-processing filters provide about 45% and 49% Bjøntegaard Delta Rate gains over VVC in instance segmentation and object detection tasks, respectively. Jukka I. Ahonen, Ramin Ghaznavi Youvalari, Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Miska M. Hannuksela, Esa Rahtu |
ISM | 2 |
| 2021 | Content-adaptive convolutional neural network post-processing filterabstractNeural Network (NN)-based coding techniques are being developed for hybrid video coding schemes, such as the Versatile Video Coding (VVC) standard. In-loop filters and postprocessing filters are two types of coding tools that aim to improve the visual quality of the reconstructed content. These tools are usually trained on large video or image datasets with varying content, but they are rarely adaptive to different content types. This problem is addressed with the proposed content-adaptive Convolutional Neural Network (CNN) post-processing filter. The proposed approach is content-adaptive in two ways. Firstly, a relatively simple CNN is pre-trained on a general video dataset and then fine-tuned on the video to be coded. Since only the bias terms of the CNN are fine-tuned, the signalling overhead is reduced. Secondly, a scaling factor indicates the influence of the CNN post-processing filter on the final reconstruction. The CNN post-processing filter is evaluated on top of VVC Test Model (VTM) 11.0 with NN-based Video Coding (NNVC) 1.0 and, overall, it can save 2.37% (Y), 3.63% (U), 2.24% (V) Bjøntegaard Delta rate (BD-rate) in the Random Access (RA) configuration. María Santamaría 0001, Yat-Hong Lam, Francesco Cricri, Jani Lainema, Ramin Ghaznavi Youvalari, Honglei Zhang 0001, Miska M. Hannuksela, Esa Rahtu, Moncef Gabbouj |
ISM | 5 |
| 2021 | Enhancing Image Coding for Machines with Compressed Feature ResidualsabstractAs computer vision technologies have tremendously improved over the last decade, videos and images are often consumed by machines instead of humans which are the main target for traditional video codecs. In many use cases, although machines are the main consumers, human involvement is also required, or even mandatory. In this paper, we propose a novel image coding technique targeted for machines, while maintaining the capability for human consumption. Our proposed codec generates two bitstreams: one bitstream from a traditional codec, referred to as human bitstream, optimized for human consumption; the other bitstream, referred to as machine bitstream, generated from an end-to-end learned neural network-based codec and optimized for machine tasks. Instead of working on the image domain, the proposed machine bitstream is derived from feature residuals – the difference between the features extracted from the input image and the features extracted from the reconstructed image generated by the traditional codec. With the help of the machine bitstream, we can significantly improve machine task performance in the low bitrate range. Our system beats the state-of-the-art traditional codec, the Versatile Video Coding (VVC/H.266), achieving −40.5% in Bjontegaard delta bitrate reduction on average for bitrates up to 0.07 BPP. Joni Seppälä, Honglei Zhang 0001, Nam Le 0003, Ramin Ghaznavi Youvalari, Francesco Cricri, Hamed Rezazadegan Tavakoli, Emre Aksu, Miska M. Hannuksela, Esa Rahtu |
ISM | 4 |
| 2021 | Adaptation and Attention for Neural Video CodingabstractNeural image coding represents now the state-of-the-art image compression approach. However, a lot of work is still to be done in the video domain. In this work, we propose an end-to-end learned video codec that introduces several architectural novelties as well as training novelties, revolving around the concepts of adaptation and attention. Our codec is organized as an intra-frame codec paired with an inter-frame codec. As one architectural novelty, we propose to train the inter-frame codec model to adapt the motion estimation process based on the resolution of the input video. A second architectural novelty is a new neural block that combines concepts from split-attention based neural networks and from DenseNets. Finally, we propose to overfit a set of decoder-side multiplicative parameters at inference time. Through ablation studies and comparisons to prior art, we show the benefits of our proposed techniques in terms of coding gains. We compare our codec to VVC/H.266 and RLVC, which represent the state-of-the-art traditional and end-to-end learned codecs, respectively, and to the top performing end-to-end learned approach in 2021 CLIC competition, E2E_T_OL. Our codec clearly outperforms E2E_T_OL, and compare favorably to VVC and RLVC in some settings. Nannan Zou, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Hamed Rezazadegan Tavakoli, Jani Lainema, Emre Aksu, Miska M. Hannuksela, Esa Rahtu |
ISM | 4 |
| 2021 | Regression-Based Motion Vector Field for Video CodingabstractIn this paper, we study a method for compensating the non-translational motion behavior in video coding. The proposed method models the motion field of a prediction block based on the motion information of the neighboring blocks by using a linear regression approach. In order to provide a finer granularity of motion vectors the Regression-based Motion Vector Field (RMVF) method derives the motion field in 4×4 sub-block accuracy. Such approach generates a smooth and more realistic motion vector field inside the prediction block. The motion field generated with RMVF is then used as a new merge mode along with other merge modes in VTM-2.0 test model of the Versatile Video Coding (H.266/VVC) standard. The conducted experiments with JVET CTC sequences illustrate that the proposed RMVF method provides 0.77%, 0.19% and 0.41% bitrate reductions with random access (RA), low delay B (LDB) and low delay P (LDP) configurations, respectively. Furthermore, this method provides on average 0.65% bitrate saving for the 360° sequences in equirectangular projection format (ERP) with RA configuration. Ramin Ghaznavi Youvalari, Alireza Aminlou, Jani Lainema |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Linear Model-Based Intra Prediction in VVC Test ModelabstractThis paper studies a new intra prediction method based on a linear model for improving the intra prediction performance of Versatile Video Coding (H.266/VVC) standard. The Linear Model-based Intra Prediction (LMIP) method in this work attempts to model the samples behavior of a coding block based on the reconstructed pixels in the neighboring of that block. The proposed method uses a 3-parameter linear function as prediction model in which the parameters of the model are derived based on a linear regression with mean square error minimization approach from the neighboring samples and their locations. The proposed LMIP method is then used as a new intra prediction mode in the VTM-4.0 test model of the VVC standard. The conducted experiments illustrate that the LMIP method provides on average 0.30% and 0.14% BD-rate improvements in luma component with all intra (AI) and random access (RA) configurations, respectively. Ramin Ghaznavi Youvalari |
ICASSP | 1 |
| 2020 | Joint Cross-Component Linear Model For Chroma Intra PredictionabstractThe Cross-Component Linear Model (CCLM) is an intra prediction technique that is adopted into the upcoming Versatile Video Coding (VVC) standard. CCLM attempts to reduce the inter-channel correlation by using a linear model. For that, the parameters of the model are calculated based on the reconstructed samples in luma channel as well as neighboring samples of the chroma coding block. In this paper, we propose a new method, called as Joint Cross-Component Linear Model (J-CCLM), in order to improve the prediction efficiency of the tool. The proposed J-CCLM technique predicts the samples of the coding block with a multi-hypothesis approach which consists of combining two intra prediction modes. To that end, the final prediction of the block is achieved by combining the conventional CCLM mode with an angular mode that is derived from the co-located luma block. The conducted experiments in VTM-8.0 test model of VVC illustrated that the proposed method provides on average more than 1.0% BD-Rate gain in chroma channels. Furthermore, the weighted YCbCr bitrate savings of 0.24% and 0.54% are achieved in 4:2:0 and 4:4:4 color formats, respectively. Ramin Ghaznavi Youvalari, Jani Lainema |
MMSP | 1 |
| 2019 | Shared Coded Picture Technique for Tile-Based Viewport-Adaptive Streaming of Omnidirectional VideoabstractTile-based viewport-adaptive streaming methods have been used in delivering omnidirectional video for virtual reality applications. In these methods, the 360° video is encoded in multiple quality versions by using the motion constrained tile set (MCTS) technique. A set of high-quality and low-quality tiles, corresponding to viewport and non-viewport areas, respectively, are selected and transmitted to the user. However, these methods require frequent intra random access points to ensure seamless viewport switching capability, very high decoding complexity, or a multi-layer coding scheme. The frequent intra random access points include very high bitrate in viewport switching points. The high decoding complexity and multi-layer decoder requirements are not aligned with the omnidirectional media format (OMAF) standard. Such requirements make these methods sub-optimal or impractical for streaming the omnidirectional video. This paper studies the current tile-based solutions for delivering the omnidirectional content. Moreover, the OMAF-compliant shared coded picture (SCP)-based scheme is proposed in this paper for streaming the omnidirectional video. The core concept of the SCP-based method is to manipulate the switching point pictures in a way that the frequent intra-coded pictures are no longer required for the viewport switching operations between different quality versions of the content. The experiments illustrated that the SCP-based method outperforms the MCTS-based method on average by 11% to 14% in terms of streaming bitrate reduction with only 4% extra decoding complexity. Ramin Ghaznavi Youvalari, Alireza Zare, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Geometry-Based Motion Vector Scaling for Omnidirectional Video CodingabstractVirtual reality (VR) applications make use of 360° omnidirectional video content for creating immersive experience to the user. In order to utilize current 2D video compression standards, such content must be projected onto a 2D image plane. However, the projection from spherical to 2D domain introduces deformations in the projected content due to the different sampling characteristics of the 2D plane. Such deformations are not favorable for the motion models of the current video coding standards. Consequently, omnidirectional video is not efficiently compressible with current codecs. In this work, a geometry-based motion vector scaling method is proposed in order to compress the motion information of omnidirectional content efficiently. The proposed method applies a scaling technique, based on the location in the 360° video, to the motion information of the neighboring blocks in order to provide a uniform motion behavior in a certain part of the content. The uniform motion behavior provides optimal candidates for efficiently predicting the motion vectors of the current block. The conducted experiments illustrated that the proposed method provides up to 2.2% bitrate reduction and on average around 1% bitrate reduction for the content with high motion characteristics in the VTM test model of Versatile Video Coding (H.266/VVC) standard. Ramin Ghaznavi Youvalari, Alireza Aminlou |
ISM | 1 |
| 2018 | Adaptive Motion Vector Prediction for Omnidirectional VideoabstractOmnidirectional video is widely used in virtual reality applications in order to create the immersive experience to the user. Such content is projected onto a 2D image plane in order to make it suitable for compression purposes by using current standard codecs. However, the resulted projected video contains deformations mainly due to the oversampling of the projection plane. These deformations are not favorable for the motion models that are used in the recent video compression standards. Hence, omnidirectional video is not efficiently compressible with the current codecs. In this work, an adaptive motion vector prediction method is proposed for efficiently coding the motion information of such content. The proposed method adaptively models the motion vectors of the coding block based on the motion information of the neighboring blocks and calculates a more optimal motion vector predictor for coding the motion information. The experimented results showed that the proposed motion vector prediction method provides up to 2.2% bitrate reduction in the content with high motion and on average 1.1% bitrate reduction for the tested sequences. Ramin Ghaznavi Youvalari, Alireza Aminlou |
VCIP | 1 |
| 2017 | Comparison of HEVC coding schemes for tile-based viewport-adaptive streaming of omnidirectional videoabstractVirtual reality applications make use of 360-degree panoramic or omnidirectional video with high resolution and high frame rate in order to create the immersive experience to the user. The user views only a portion of the captured 360-degree scene at each time instant, hence streaming the whole omnidirectional video in highest quality is not efficient. In order to alleviate the problem of bandwidth wastage, viewport-adaptive encoding and streaming schemes have been proposed. In these schemes, part of the captured scene that is within the viewer's field of view is delivered at highest quality while the rest of the scene in a lower quality. In this work, three tile-based viewport-adaptive methods using motion-constrained tile sets (MCTS), region-of-interest scalability and simulcast approach have been studied for streaming omnidirectional content. In the performed experiments with various tiling arrangements, MCTS-based scheme required highest bitrate compared to other methods. The scalable coding scheme provided the highest performance in terms of streaming bitrate saving on average up to 53% and 35% compared to streaming the whole omnidirectional video and MCTS-based method, respectively. Ramin Ghaznavi Youvalari, Alireza Zare, Huameng Fang, Alireza Aminlou, Qingpeng Xie, Miska M. Hannuksela, Moncef Gabbouj |
MMSP | 1 |
| 2016 | Efficient Coding of 360-Degree Pseudo-Cylindrical Panoramic Video for Virtual Reality ApplicationsabstractPseudo-cylindrical panoramas represent the data distribution of spherical coordinates closely in two-dimensional domain due to the equidistant sampling of 360-degree scene. Therefore, unlike the cylindrical projections, they do not suffer from the over stretching in the polar areas. However, due to the non-rectangular format in effective picture area and sharp edges at its borders, the compression performance is inefficient. In this paper, we propose two methods which improve the compression performance of both intra-frame and inter-frame coding of pseudo-cylindrical panoramic content and meanwhile reduce the coding artifacts. In the intra-frame coding method, border edges are smoothed by modifying the content of the image in the non-effective picture area, which are cropped at the receiver side. In the inter-frame coding method, gaining the benefit of 360-degree property of the content, non-effective picture area of reference frames at border is filled with the content of the effective picture area from the opposite border to enhance the performance of motion compensation. Ramin Ghaznavi Youvalari, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj |
ISM | 1 |
| 2016 | Analysis of regional down-sampling methods for coding of omnidirectional videoabstractIn order to compress omnidirectional video clips, a projection onto a two-dimensional image plane is necessary. The most commonly used projection format is the equirectangular panoramic projection, which results into a significant amount of redundant samples in the polar areas. The redundant samples incur extra bitrate and increase the encoding/decoding time. In this paper, we study regional down-sampling (RDS) for achieving better compression and smaller encoding/decoding time for omnidirectional content. We extend the persistent RDS method applied equally to all pictures to be applied to selected pictures only in our proposed temporal RDS method and then compare the persistent and temporal RDS methods. The simulation results indicate that both the persistent and temporal RDS improve the rate-distortion (RD) performance compared to the conventional coding of equirectangular panoramas, while the temporal RDS method has less sequence-wise RD performance variation and slightly better RD performance on average when compared to the persistent RDS technique. Alongside the coding methods, we study spherical quality measurement methods for VR images/video and analyze the coding methods with these quality metrics. Moreover, we propose a uniformly sampled spherical quality metric in order to evaluate the coding distortion of omnidirectional videos. Ramin Ghaznavi Youvalari, Alireza Aminlou, Miska M. Hannuksela |
PCS | 1 |