VLDB 2026 Research / reviewers in the wild / expert
Francesco Cricri
dblp:95/7814
· DBLP profile ↗
42ranked-venue papers
7as first author
26since 2021 · last 2025
0000-0002-1521-420XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 6 first-author · 26 since 2021Artificial intelligence and machine learning · 6 · 1 since 2021Computer networks · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learned Image Codec with Progressive Multi-Scale Probability Model for Streaming in Unreliable Communication ChannelsabstractVideo streaming over the Internet is one of the most important applications of video compression technologies. During streaming, network congestion and packet loss can catastrophically corrupt conventional block-based codecs, yielding fully corrupted frames or partially visible content that significantly degrades user experience. Unlike block-based conventional codecs, end-to-end learned codecs operate on holistic feature representations, unlocking a new paradigm for error resilience. In this work, we leverage a progressive multi-scale entropy model to partition the bitstream into ordered data units, ensuring that any received prefix unit yields a low-quality but full-frame reconstruction. To handle missing information in the latent tensor, we introduce lightweight adapters that predict absent features before image synthesis. Experiments on JVET CTC classes show that our methods improve PSNR by up to 2.27 dB over maximum-likelihood tensor filling, at no extra bitrate cost. Honglei Zhang 0001, A. Burakhan Koyuncu, Jukka I. Ahonen, Nannan Zou, Francesco Cricri |
MMSP | 5 |
| 2025 | Task Enhancement Tiles for Ultra Lightweight Post-processing in Visual Coding for MachinesabstractThe proliferation of automated visual analysis calls for compression methods tailored to the unique requirements of Video Coding for Machines (VCM). In this paper, we propose a computationally lightweight post-processing method that is based on a learned component referred to as a task enhancement tile (TET). A TET is spatially tiled over the reconstructed visual data and added to it element-wise. It only requires one addition per pixel in each color channel before the machine task can be applied. Our results with the VVC test model (VTM) demonstrate coding gains of up to 39.0% for object detection and 29.2% for instance segmentation on image datasets, while evaluation on a video dataset shows gains of up to 35.2% for object detection, relative to the VTM anchor. The proposed solution also offers extremely low computational cost, preservation of human-viewable content, full compliance with video coding standards, no requirement for side information transmission from encoder to decoder, and generalization across tasks, models, and encoding parameters. Tero Partanen, Alban Marie, Rudolf Kortelahti, Alexandre Mercat, Jarno Vanne, Miska M. Hannuksela, Honglei Zhang 0001, Alireza Aminlou, Francesco Cricri |
PCS | 9 |
| 2025 | A Hybrid Framework Integrating End-to-End Learned Image Codec with Conventional Codec
Nannan Zou, Antti Hallapuro, Francesco Cricri, Honglei Zhang 0001, A. Burakhan Koyuncu, Jukka I. Ahonen, Miska M. Hannuksela, Esa Rahtu |
PCS | 3 |
| 2024 | Competitive Learning For Achieving Content-Specific Filters In Video Coding For MachinesabstractThis paper investigates the efficacy of jointly optimizing content-specific post-processing filters to adapt a human-oriented video/image codec into a codec suitable for machine vision tasks. By observing that artifacts produced by video/image codecs are content-dependent, we propose a novel training strategy based on competitive learning principles. This strategy assigns training samples to filters dynamically, in a fuzzy manner, which further optimizes the winning filter on the given sample. Inspired by simulated annealing optimization techniques, we employ a softmax function with a temperature variable as the weight allocation function to mitigate the effects of random initialization. Our evaluation, conducted on a system utilizing multiple post-processing filters within a Versatile Video Coding (VVC) codec framework, demonstrates the superiority of content-specific filters trained with our proposed strategies, specifically, when images are processed in blocks. Using VVC reference software VTM 12.0 as the anchor, experiments on the OpenImages dataset show an improvement in the BD-rate reduction from -41.3% and -44.6% to -42.3% and -44.7% for object detection and instance segmentation tasks, respectively, compared to independently trained filters. The statistics of the filter usage align with our hypothesis and underscore the importance of jointly optimizing filters for both content and reconstruction quality. Our findings pave the way for further improving the performance of video/image codecs. Honglei Zhang 0001, Jukka I. Ahonen, Nam Le 0003, Ruiying Yang, Francesco Cricri |
ICIP | 5 |
| 2024 | Feasibility Study of Multi-Layer VVC Coding Scheme for Hybrid Machine-Human ConsumptionabstractThe proliferation of machine vision applications necessitates developing more efficient visual data compression schemes for machine consumption. However, numerous automated use cases still require keeping humans in the loop, leading to the need for a machine-optimized video streaming with the option for human supervision. This paper investigates the feasibility of using the multi-layer coding approach of the emerging Versatile Video Coding (VVC) standard to create favorable conditions for hybrid machine-human consumption. We introduce a multi-layer coding scheme, where the base layer (BL) is optimized for machines and the enhancement layer (EL) complements the stream for human vision. Our results demonstrate that the bitrate of the proposed multi-layer stream (BL + EL) is, on average, 11% higher than that of a single-layer VVC. However, the more compact BL yields overall bandwidth savings as long as the EL is required less than 80% of the time. Jaakko Laitinen, Tero Partanen, Alexandre Mercat, Jarno Vanne, Miska M. Hannuksela, Honglei Zhang 0001, Alireza Aminlou, Francesco Cricri |
ICME | 8 |
| 2024 | Luma Range Scaling for Enhanced VVC Efficiency in Video Coding for MachinesabstractRecent years have shown significant growth in video data traffic for machine vision applications, catalyzing new standardization efforts in video coding for machines (VCM). These activities focus on compressing images and videos for machine vision tasks, rather than for human viewing. In this work, we propose a novel method that scales down the luma range to enhance the coding efficiency of Versatile Video Coding (VVC) for machine consumption. This method results in a lower bitrate after encoding and has only minimal adverse effects on the accuracy of machine vision tasks. In our experiments, we down-scale the luma channel of the input video using luma-scaling factors from 0.2 to 0.9 and evaluate coding results with optional back-scaling to the original range before machine vision tasks. Our results with the VVC Test Model (VTM) demonstrate that the proposed technique achieves coding gain of up to 37.9%and 46.1% for the same object detection and tracking accuracy, respectively. Tero Partanen, Alban Marie, Alexandre Mercat, Jarno Vanne, Miska M. Hannuksela, Honglei Zhang 0001, Alireza Aminlou, Francesco Cricri |
MMSP | 8 |
| 2023 | NN-VVC: Versatile Video Coding boosted by self-supervisedly learned image coding for machinesabstractThe recent progress in artificial intelligence has led to an ever-increasing usage of images and videos by machine analysis algorithms, mainly neural networks. Nonetheless, compression, storage and transmission of media have traditionally been designed considering human beings as the viewers of the content. Recent research on image and video coding for machine analysis has progressed mainly in two almost orthogonal directions. The first is represented by end-to-end (E2E) learned codecs which, while offering high performance on image coding, are not yet on par with state-of-the-art conventional video codecs and lack interoperability. The second direction considers using the Versatile Video Coding (VVC) standard or any other conventional video codec (CVC) together with pre- and post-processing operations targeting machine analysis. While the CVC-based methods benefit from interoperability and broad hardware and software support, the machine task performance is often lower than the desired level, particularly in low bitrates. This paper proposes a hybrid codec for machines called NN-VVC, which combines the advantages of an E2E-learned image codec and a CVC to achieve high performance in both image and video coding for machines. Our experiments show that the proposed system achieved up to -43.20% and -26.8% Bjøntegaard Delta rate reduction over VVC for image and video data, respectively, when evaluated on multiple different datasets and machine vision tasks. To the best of our knowledge, this is the first research paper showing a hybrid video codec that outperforms VVC on multiple datasets and multiple machine vision tasks. Jukka I. Ahonen, Nam Le 0003, Honglei Zhang 0001, Antti Hallapuro, Francesco Cricri, Hamed Rezazadegan Tavakoli, Miska M. Hannuksela, Esa Rahtu |
ISM | 5 |
| 2023 | Region of Interest Enabled Learned Image Coding for MachinesabstractImage and video coding for machines has been recently gaining more and more interest from both the industry and the research community. One successful approach is based on end-to-end (E2E) learned compression and has shown significant gains over the state-of-the-art conventional image coding methods. However, one of the remaining challenges for such E2E-learned image codecs for machines is to adaptively allocate the bits over different regions of the image, while retaining the machine vision performance. In this paper, we propose a method that leverages Regions-Of-Interest (ROIs) for bitrate allocation within a Learned Image Codec (LIC) for machines. In particular, the proposed method reduces the bits allocated for the background regions of the image by reducing the variance of the elements corresponding to the background regions in the latent representation. This results in more heavily quantized background areas, while keeping the quality of the ROI areas suitable for machine tasks. The proposed method achieves significant gains, -15.80% and -22.43% Pareto BD-rate reduction, over the baseline LIC on object detection and instance segmentation tasks, respectively. To the best of our knowledge, this is the first research paper proposing an ROI-based inference-time technology for Learned Image Coding for machines. Jukka I. Ahonen, Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Esa Rahtu |
MMSP | 4 |
| 2023 | Overfitting NN loop-filters in video codingabstractOverfitting is usually regarded as a negative condition since it impairs the generalisation power of a model. Nevertheless, overfitting a Neural Network (NN) on test data may be advantageous to improve the compression efficiency of image/video coding tools and systems. Previous research has demonstrated the benefits of NN overfitting for post-processing operations, i.e. post-filters, but not yet for actual decoding tools. Generally, the NN is overfitted on test data at the encoder end, and the weight update is coded and sent to the decoder end along the image/video bitstream. The proposed approach follows this strategy. In particular, the overfitting of the Low Operation Point (LOP) loop-filter in NN-based Video Coding (NNVC) software is studied. The overall approach yields Bjøntegaard Delta rate (BD-rate) of -7.74%, -13.73% and -12.49%, for the Y, U and V components, respectively. Out of these coding gains, 1.21%, 6.43% and 5.52%, for the Y, U and V components, are attributed to the overfitting. The boost in the coding gains comes with only 1.5% more complexity, due to the multiplier parameters introduced during the overfitting. Ruiying Yang, María Santamaría 0001, Francesco Cricri, Honglei Zhang 0001, Jani Lainema, Ramin Ghaznavi Youvalari, Miska M. Hannuksela, Tapio Elomaa |
VCIP | 3 |
| 2022 | Content-Adaptive Neural Network Post-Processing Filter with NNR-Coded Weight-UpdatesabstractNeural Network (NN) filters improve the perceptual quality of reconstructed videos by reducing compression artefacts. For content adaptation, a few NN-filters use over-fitting. As the adaptation signal is a weight-update, compression is required to minimise significant bitrate overheads. Most approaches, however, use generic data compression algorithms, which are inadequate for coding NN weight-updates. This work introduces a content-adaptive NN post-processing filter with weight-updates coded using the Neural Network compression and Representation (NNR) standard. The bitrate overhead is further decreased by over-fitting only a subset of weights, selected via energy-based analysis. The proposed filter saved about 4.57% (Y), 10.33% (Cb), 6.53% (Cr) Bjøntegaard Delta rate (BD-rate) on top of the Versatile Video Coding (VVC) Test Model (VTM) 11.0 with NN-based Video Coding (NNVC) 1.0, in Random Access (RA) configuration. Compared to the non-over-fitted NN, the performance was doubled; and compared to 7z, NNR reduced the bitrate of the weight-update by ∼64%. María Santamaría 0001, Francesco Cricri, Jani Lainema, Ramin Ghaznavi Youvalari, Honglei Zhang 0001, Miska M. Hannuksela |
ICIP | 2 |
| 2022 | Bridging the Gap Between Image Coding for Machines and HumansabstractImage coding for machines (ICM) aims at reducing the bitrate required to represent an image while minimizing the drop in machine vision analysis accuracy. In many use cases, such as surveillance, it is also important that the visual quality is not drastically deteriorated by the compression process. Recent works on using neural network (NN) based ICM codecs have shown significant coding gains against traditional methods; however, the decompressed images, especially at low bitrates, often contain checkerboard artifacts. We propose an effective decoder finetuning scheme based on adversarial training to significantly enhance the visual quality of ICM codecs, while preserving the machine analysis accuracy, without adding extra bitcost or parameters at the inference phase. The results show complete removal of the checkerboard artifacts at the negligible cost of −1.6% relative change in task performance score. In the cases where some amount of artifacts is tolerable, such as when machine consumption is the primary target, this technique can enhance both pixel-fidelity and feature-fidelity scores without losing task performance. Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Hamed Rezazadegan Tavakoli, Emre Aksu, Miska M. Hannuksela, Esa Rahtu |
ICIP | 3 |
| 2022 | Stochastic Binary-Ternary Quantization for Communication Efficient Federated ComputationabstractA stochastic binary-ternary (SBT) quantization approach is introduced for communication efficient federated computation; form of collaborative computing where locally trained models are exchanged between institutes. Communication of deep neural network models could be highly inefficient due to their large size. This motivates model compression in which quantization is an important step. Two well-known quantization algorithms are binary and ternary quantization. The first leads into good compression, sacrificing accuracy. The second provides good accuracy with less compression. To better benefit from trade-off between accuracy and compression, we propose an algorithm to stochastically switch between binary and ternary quantization. By combining with uniform quantization, we further extend the proposed algorithm to a hierarchical method which results in even better compression without sacrificing the accuracy. We tested the proposed algorithm using Neural network Compression Test Model (NCTM) provided by MPEG community. Our results demonstrate that the hierarchical variant of the proposed algorithm outperforms other quantization algorithms in term of compression, while maintaining the accuracy competitive to that provided by other methods. Rangu Goutham, Homayun Afrabandpey, Francesco Cricri, Honglei Zhang 0001, Emre Aksu, Miska M. Hannuksela, Hamed Rezazadegan Tavakoli |
ICIP | 3 |
| 2022 | Adaptive Multi-Scale Progressive Probability Model for Lossless Image CompressionabstractDomain adaptation is an efficient technique to improve the performance of a system by adapting a pre-trained model to the given input data. The adaptation technique has been generally applied in conventional video codecs. For neural network-based systems, the encoder may adapt the decoder to the input data by fine-tuning a pre-trained model present at the decoder side. The weight update is then transferred to the decoder and the updated model is used to decode the bitstream. However, due to the large number of parameters in deep neural networks, the overhead of the weight update may diminish the gain from the adaptation technique. In recent years, various methods have been proposed to reduce the overhead without significantly compromising the gain. In this paper, we propose an adaptive multi-scale progressive probability model for lossless image compression. The proposed method uses the data that has already been processed at the inference stage to fine-tune the probability model. Importantly, the decoder can apply the fine-tuning by itself resulting a small adaptation overhead to help the decoder in performing the fine-tuning. The proposed method achieves up to 0.28 bits-per-pixel (BPP) reduction on four benchmark datasets compared to the state-of-the-art method. Honglei Zhang 0001, Francesco Cricri, Nannan Zou, Hamed Rezazadegan Tavakoli, Miska M. Hannuksela |
ICIP | 2 |
| 2022 | TMD: Transformed Mesh Decoder for Mesh AnimationabstractEasy and fast animation of 3D characters is attractive for both gaming and entertainment applications. 3D mesh must be properly rigged and skinned in order to create seamless animation. This process can be time consuming and requires deep knowledge of appropriate software based on kinematic animation. In this work, we present a fast and lightweight deep neural model to automate 3D human animation using skeletal representation from 2D image pose, i.e., joints in 2D space. We accomplish this using Transformed Mesh Decoder (TMD), which is a novel layer for convolutional neural networks. To train the network, we generate a large and diverse dataset using Skinned Multi-Person Linear (SMPL) model. Experiment shows that our method is effective when compared to both the ground truth and state-of-the art linear blend skinning that require manually painted skinning weights for accurate result. The animation process is fast and can achieve approximately 10-15fps in practice. The proposed method is simple which opens the possibility for future improvement in real-time application. Peter Fasogbon, Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Emre Aksu |
ICPR | 3 |
| 2022 | Low-precision post-filtering in video codingabstractNeural Networks (NNs) have demonstrated their effectiveness in tackling challenges involving multimedia content. In the video coding field, NNs are actively exploited as novel tools that complement conventional signal processing tools, as well as end-to-end coding solutions. Since NNs use commonly floating-point arithmetic, different results may be generated in different computing environments, leading to discrepancies and even corrupted reconstructions. Accordingly, this issue is solved by employing fixed-point arithmetic instead. This paper studies the quantisation of a 32-bit floating-point (float32) NN post-filter to 32-bit fixed-point (int32) and 16-bit fixed-point (int16). On top of the VVC Test Model (VTM) 11.0 NN-based Video Coding (NNVC) 1.0, the coding gains of the float32 post-filter are 5.01% (Y), 18.95% (Cb) and 17.33% Cr. Compared to the float32 inference, the quantised models produce coding losses: 0.01% (Y, Cb and Cr) for the int32 inference and 0.45% (Y), 1.97% (Cb) and 1.19% (Cr) for the int16 inference. Nevertheless, the fixed-point approaches achieve bit exact matches in different computing environments. Moreover, the decoding time with int16 is about half the decoding time of int32. Ruiying Yang, María Santamaría 0001, Francesco Cricri, Honglei Zhang 0001, Jani Lainema, Ramin Ghaznavi Youvalari, Miska M. Hannuksela |
ISM | 3 |
| 2022 | The Lottery Ticket Adaptation for Neural Video CodingabstractRecently, learning based video compression methods have attracted increasing attention. However, most learning based video codecs are not adaptive to different video contents. Though adaptation at inference time is a solution to tackle this issue, adapting all the codec’s parameters is computationally expensive and brings heavy bitrate overhead. The recently proposed Lottery Ticket Hypothesis (LTH) states that an over-parameterized neural network contains smaller subnetworks (winning tickets) that can match the performance of the original network. In this paper, we present a novel lottery-ticket adaptation technique on decoder-side multiplicative parameters of a neural network, transferring the concept of winning lottery tickets to video compression tasks. At inference time, the winning multiplicative parameters are overfitted, compressed, and signaled together with encoded frames for decoding. We show that our approach outperforms the Versatile Video Coding (VVC) standard in the Multiscale Structural Similarity (MS-SSIM) at a low bitrate on both the UVG and JVET sequences. To the best of our knowledge, this is the first attempt to apply LTH in the video compression domain. Also, this is the first published end-to-end learned video codec working directly on YUV format, which outperforms VVC on UVG and JVET datasets in MS-SSIM. Nannan Zou, Francesco Cricri, Honglei Zhang 0001, Hamed Rezazadegan Tavakoli, Miska M. Hannuksela, Esa Rahtu |
ISM | 2 |
| 2022 | Overview of the Neural Network Compression and Representation (NNR) StandardabstractNeural Network Coding and Representation (NNR) is the first international standard for efficient compression of neural networks (NNs). The standard is designed as a toolbox of compression methods, which can be used to create coding pipelines. It can be either used as an independent coding framework (with its own bitstream format) or together with external neural network formats and frameworks. For providing the highest degree of flexibility, the network compression methods operate per parameter tensor in order to always ensure proper decoding, even if no structure information is provided. The NNR standard contains compression-efficient quantization and deep context-adaptive binary arithmetic coding (DeepCABAC) as core encoding and decoding technologies, as well as neural network parameter pre-processing methods like sparsification, pruning, low-rank decomposition, unification, local scaling and batch norm folding. NNR achieves a compression efficiency of more than 97% for transparent coding cases, i.e. without degrading classification quality, such as top-1 or top-5 accuracies. This paper provides an overview of the technical features and characteristics of NNR. Heiner Kirchhoffer, Paul Haase, Wojciech Samek, Karsten Müller 0001, Hamed Rezazadegan Tavakoli, Francesco Cricri, Emre Aksu, Miska M. Hannuksela, Wei Jiang 0001, Wei Wang 0311, Shan Liu 0001, Swayambhoo Jain, Shahab Hamidi-Rad, Fabien Racapé, Werner Bailer |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | Image Coding For Machines: an End-To-End Learned ApproachabstractOver recent years, deep learning-based computer vision systems have been applied to images at an ever-increasing pace, oftentimes representing the only type of consumption for those images. Given the dramatic explosion in the number of images generated per day, a question arises: how much better would an image codec targeting machine-consumption perform against state-of-the-art codecs targeting human-consumption? In this paper, we propose an image codec for machines which is neural network (NN) based and end-to-end learned. In particular, we propose a set of training strategies that address the delicate problem of balancing competing loss functions, such as computer vision task losses, image distortion losses, and rate loss. Our experimental results show that our NN-based codec outperforms the state-of-the-art Versa-tile Video Coding (VVC) standard on the object detection and instance segmentation tasks, achieving -37.87% and -32.90% of BD-rate gain, respectively, while being fast thanks to its compact size. To the best of our knowledge, this is the first end-to-end learned machine-targeted image codec. Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Esa Rahtu |
ICASSP | 3 |
| 2021 | Mind The Structure: Adopting Structural Information For Deep Neural Network CompressionabstractDeep neural networks have huge number of parameters and require large number of bits for representation. This hinders their adoption in decentralized environments where model transfer among different parties is a characteristic of the environment while the communication bandwidth is limited. Parameter quantization is a compression approach to address this challenge by reducing the number of bits required to represent a model, e.g. a neural network. However, majority of existing neural network quantization methods do not exploit structural information of layers and parameters during quantization. In this paper, focusing on Convolutional Neural Networks (CNNs), we present a novel quantization approach by employing the structural information of neural network layers and their corresponding parameters. Starting from a pre-trained CNN, we categorize network parameters into different groups based on the similarity of their layers and their spatial structure. Parameters of each group are independently clustered and the centroid of each cluster is used as representative for all parameters in the cluster. Finally, the centroids and the cluster indexes of the parameters are used as a compact representation of the parameters. Experiments with two different tasks, i.e., acoustic scene classification and image compression, demonstrate the effectiveness of the proposed approach. Homayun Afrabandpey, Anton Muravev, Hamed Rezazadegan Tavakoli, Honglei Zhang 0001, Francesco Cricri, Moncef Gabbouj, Emre Aksu |
ICIP | 5 |
| 2021 | Hybrid Pruning And SparsificationabstractA hybrid approach based on the combination of saliency-based neural pruning and regularization-based sparsification is proposed. We propose using a graph diffusion process for determining the neuron importance for pruning. Then, we use a regularization loss based on weighted $L_{1}-$norm and $L_{2}-$norm during fine-tuning to recover the lost performance. This is followed by a threshold step to further impose sparsification. We demonstrate such a hybrid approach achieves significantly better performance in comparison to purely regularization-based sparsification for large neural networks. To this end, we assessed our proposed method on three tasks, including: image classification (3 network architectures), audio classification and image compression. Hamed Rezazadegan Tavakoli, Joachim Wabnig, Francesco Cricri, Honglei Zhang 0001, Emre Aksu, Iraj Saniee |
ICIP | 3 |
| 2021 | Learned Image Coding for Machines: A Content-Adaptive ApproachabstractToday, according to the Cisco Annual Internet Report (2018-2023), the fastest-growing category of Internet traffic is machine-to-machine communication. In particular, machine-to-machine communication of images and videos represents a new challenge and opens up new perspectives in the context of data compression. One possible solution approach consists of adapting current human-targeted image and video coding standards to the use case of machine consumption. Another approach consists of developing completely new compression paradigms and architectures for machine-to-machine communications. In this paper, we focus on image compression and present an inference-time content-adaptive fine-tuning scheme that optimizes the latent representation of an end-to-end learned image codec, aimed at improving the compression efficiency for machine-consumption. The conducted experiments targeting instance segmentation task network show that our online finetuning brings an average bitrate saving (BD-rate) of -3.66% with respect to our pretrained image codec. In particular, at low bitrate points, our proposed method results in a significant bitrate saving of -9.85%. Overall, our pretrained-and-then-finetuned system achieves - 30.54% BD-rate over the state-of-the-art image/video codec Versatile Video Coding (VVC) on instance segmentation. Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Hamed Rezazadegan Tavakoli, Esa Rahtu |
ICME | 3 |
| 2021 | Learned Enhancement Filters for Image Coding for MachinesabstractMachine-To-Machine (M2M) communication applications and use cases, such as object detection and instance segmentation, are becoming mainstream nowadays. As a consequence, majority of multimedia content is likely to be consumed by machines in the coming years. This opens up new challenges on efficient compression of this type of data. Two main directions are being explored in the literature, one being based on existing traditional codecs, such as the Versatile Video Coding (VVC) standard, that are optimized for human-targeted use cases, and another based on end-to-end trained neural networks. However, traditional codecs have significant benefits in terms of interoperability, real-time decoding, and availability of hardware implementations over end-to-end learned codecs. Therefore, in this paper, we propose learned post-processing filters that are targeted for enhancing the performance of machine vision tasks for images reconstructed by the VVC codec. The proposed enhancement filters provide significant improvements on the target tasks compared to VVC coded images. The conducted experiments show that the proposed post-processing filters provide about 45% and 49% Bjøntegaard Delta Rate gains over VVC in instance segmentation and object detection tasks, respectively. Jukka I. Ahonen, Ramin Ghaznavi Youvalari, Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Miska M. Hannuksela, Esa Rahtu |
ISM | 5 |
| 2021 | Content-adaptive convolutional neural network post-processing filterabstractNeural Network (NN)-based coding techniques are being developed for hybrid video coding schemes, such as the Versatile Video Coding (VVC) standard. In-loop filters and postprocessing filters are two types of coding tools that aim to improve the visual quality of the reconstructed content. These tools are usually trained on large video or image datasets with varying content, but they are rarely adaptive to different content types. This problem is addressed with the proposed content-adaptive Convolutional Neural Network (CNN) post-processing filter. The proposed approach is content-adaptive in two ways. Firstly, a relatively simple CNN is pre-trained on a general video dataset and then fine-tuned on the video to be coded. Since only the bias terms of the CNN are fine-tuned, the signalling overhead is reduced. Secondly, a scaling factor indicates the influence of the CNN post-processing filter on the final reconstruction. The CNN post-processing filter is evaluated on top of VVC Test Model (VTM) 11.0 with NN-based Video Coding (NNVC) 1.0 and, overall, it can save 2.37% (Y), 3.63% (U), 2.24% (V) Bjøntegaard Delta rate (BD-rate) in the Random Access (RA) configuration. María Santamaría 0001, Yat-Hong Lam, Francesco Cricri, Jani Lainema, Ramin Ghaznavi Youvalari, Honglei Zhang 0001, Miska M. Hannuksela, Esa Rahtu, Moncef Gabbouj |
ISM | 3 |
| 2021 | Enhancing Image Coding for Machines with Compressed Feature ResidualsabstractAs computer vision technologies have tremendously improved over the last decade, videos and images are often consumed by machines instead of humans which are the main target for traditional video codecs. In many use cases, although machines are the main consumers, human involvement is also required, or even mandatory. In this paper, we propose a novel image coding technique targeted for machines, while maintaining the capability for human consumption. Our proposed codec generates two bitstreams: one bitstream from a traditional codec, referred to as human bitstream, optimized for human consumption; the other bitstream, referred to as machine bitstream, generated from an end-to-end learned neural network-based codec and optimized for machine tasks. Instead of working on the image domain, the proposed machine bitstream is derived from feature residuals – the difference between the features extracted from the input image and the features extracted from the reconstructed image generated by the traditional codec. With the help of the machine bitstream, we can significantly improve machine task performance in the low bitrate range. Our system beats the state-of-the-art traditional codec, the Versatile Video Coding (VVC/H.266), achieving −40.5% in Bjontegaard delta bitrate reduction on average for bitrates up to 0.07 BPP. Joni Seppälä, Honglei Zhang 0001, Nam Le 0003, Ramin Ghaznavi Youvalari, Francesco Cricri, Hamed Rezazadegan Tavakoli, Emre Aksu, Miska M. Hannuksela, Esa Rahtu |
ISM | 5 |
| 2021 | Adaptation and Attention for Neural Video CodingabstractNeural image coding represents now the state-of-the-art image compression approach. However, a lot of work is still to be done in the video domain. In this work, we propose an end-to-end learned video codec that introduces several architectural novelties as well as training novelties, revolving around the concepts of adaptation and attention. Our codec is organized as an intra-frame codec paired with an inter-frame codec. As one architectural novelty, we propose to train the inter-frame codec model to adapt the motion estimation process based on the resolution of the input video. A second architectural novelty is a new neural block that combines concepts from split-attention based neural networks and from DenseNets. Finally, we propose to overfit a set of decoder-side multiplicative parameters at inference time. Through ablation studies and comparisons to prior art, we show the benefits of our proposed techniques in terms of coding gains. We compare our codec to VVC/H.266 and RLVC, which represent the state-of-the-art traditional and end-to-end learned codecs, respectively, and to the top performing end-to-end learned approach in 2021 CLIC competition, E2E_T_OL. Our codec clearly outperforms E2E_T_OL, and compare favorably to VVC and RLVC in some settings. Nannan Zou, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Hamed Rezazadegan Tavakoli, Jani Lainema, Emre Aksu, Miska M. Hannuksela, Esa Rahtu |
ISM | 3 |
| 2021 | Learn to overfit better: finding the important parameters for learned image compressionabstractFor most machine learning systems, overfitting is an undesired behavior. However, overfitting a model to a test image or a video at inference time is a favorable and effective technique to improve the coding efficiency of learning-based image and video codecs. At the encoding stage, one or more neural networks that are part of the codec are finetuned using the input image or video to achieve a better coding performance. The encoder en-codes the input content into a content bitstream. If the finetuned neural network is part (also) of the decoder, the encoder signals the weight update of the finetuned model to the decoder along with the content bitstream. At the decoding stage, the decoder first updates its neural network model according to the received weight update, and then proceeds with decoding the content bitstream. Since a neural network contains a large number of parameters, compressing the weight update is critical to reducing bitrate overhead. In this paper, we propose learning-based methods to find the important parameters to be overfitted, in terms of rate-distortion performance. Based on simple distribution models for variables in the weight update, we derive two objective functions. By optimizing the proposed objective functions, the importance scores of the parameters can be calculated and the important parameters can be determined. Our experiments on lossless image compression codec show that the proposed method significantly outperforms a prior-art method where overfitted parameters were selected based on heuristics. Furthermore, our technique improved the compression performance of the state-of-the-art lossless image compression codec by 0.1 bit per pixel. Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, María Santamaría 0001, Yat-Hong Lam, Miska M. Hannuksela |
VCIP | 2 |
| 2020 | Lossless Image Compression Using a Multi-scale Progressive Statistical Model
Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Nannan Zou, Emre Aksu, Miska M. Hannuksela |
ACCV (3) | 2 |
| 2020 | Efficient Adaptation of Neural Network Filter for Video CompressionabstractWe present an efficient finetuning methodology for neural-network filters which are applied as a postprocessing artifact-removal step in video coding pipelines. The fine-tuning is performed at encoder side to adapt the neural network to the specific content that is being encoded. In order to maximize the PSNR gain and minimize the bitrate overhead, we propose to finetune only the convolutional layers' biases. The proposed method achieves convergence much faster than conventional finetuning approaches, making it suitable for practical applications. The weight-update can be included into the video bitstream generatedby the existing video codecs. We show that our method achieves up to 9.7% average BD-rate gain when compared to the state-of-art Versatile Video Coding (VVC) standard codec on 7 test sequences. Yat Hong Lam, Alireza Zare, Francesco Cricri, Jani Lainema, Miska M. Hannuksela |
ACM Multimedia | 3 |
| 2020 | L2C - Learning to Learn to CompressabstractIn this paper we present an end-to-end meta-learned system for image compression. Traditional machine learning based approaches to image compression train one or more neural network for generalization performance. However, at inference time, the encoder or the latent tensor output by the encoder can be optimized for each test image. This optimization can be regarded as a form of adaptation or benevolent overfitting to the input content. In order to reduce the gap between training and inference conditions, we propose a new training paradigm for learned image compression, which is based on meta-learning. In a first phase, the neural networks are trained normally. In a second phase, the Model-Agnostic Meta-learning approach is adapted to the specific case of image compression, where the inner-loop performs latent tensor overfitting, and the outer loop updates both encoder and decoder neural networks based on the overfitting performance. Furthermore, after meta-learning, we propose to overfit and cluster the bias terms of the decoder on training image patches, so that at inference time the optimal content-specific bias terms can be selected at encoder-side. Finally, we propose a new probability model for lossless compression, which combines concepts from both multi-scale and super-resolution probability model approaches. We show the benefits of all our proposed ideas via carefully designed experiments. Nannan Zou, Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Jani Lainema, Miska M. Hannuksela, Emre Aksu, Esa Rahtu |
MMSP | 3 |
| 2019 | Approximating Binarization in Neural NetworksabstractBinarization of neural networks' activations may be a requirement for some applications. A typical example is end-to-end learned deep image compression systems where the encoder's output is requred to be a binary vector. Binarization is non-differentiable, therefore one needs to approximate it in order to train neural networks with stochastic gradient descent. In this paper, we investigate these training strategies and provide improvements over baselines. We find that during training, constraining the activations in a region that is far away from binary points leads to a better performance at test-time. The above finding provides a counter-intuitive result and leads to re-thinking the binarization approximation problem in neural networks. Çaglar Aytekin, Francesco Cricri, Jani Lainema, Emre Aksu, Miska M. Hannuksela |
IJCNN | 2 |
| 2018 | Depth Masked Discriminative Correlation FilterabstractDepth information provides a strong cue for occlusion detection and handling, but has been largely omitted in generic object tracking until recently due to lack of suitable benchmark datasets and applications. In this work, we propose a Depth Masked Discriminative Correlation Filter (DM-DCF) which adopts novel depth segmentation based occlusion detection that stops correlation filter updating and depth masking which adaptively adjusts the spatial support for correlation filter. In Princeton RGBD Tracking Benchmark, our DM-DCF is among the state-of-the-art in overall ranking and the winner on multiple categories. Moreover, since it is based on DCF, “DM-DCF” runs an order of magnitude faster than its competitors making it suitable for time constrained applications. Ugur Kart, Joni-Kristian Kämäräinen, Jiri Matas, Lixin Fan, Francesco Cricri |
ICPR | 5 |
| 2018 | Object Detection in Equirectangular PanoramaabstractWe introduce a high-resolution equirectangular panorama (aka 360-degree, virtual reality, VR) dataset for object detection and propose a multi-projection variant of the YOLO detector. The main challenges with equirectangular panorama images are i) the lack of annotated training data, ii) high-resolution imagery and iii) severe geometric distortions of objects near the panorama projection poles. In this work, we solve the challenges by I) using training examples available in the “conventional datasets” (ImageNet and COCO), II) employing only low resolution images that require only moderate GPU computing power and memory, and III) our multi-projection YOLO handles projection distortions by making multiple stereographic sub-projections. In our experiments, YOLO outperforms the other state-of-the-art detector, Faster R-CNN, and our multi-projection YOLO achieves the best accuracy with low-resolution input. Wenyan Yang, Yanlin Qian, Joni-Kristian Kämäräinen, Francesco Cricri, Lixin Fan |
ICPR | 4 |
| 2018 | Clustering and Unsupervised Anomaly Detection with l2 Normalized Deep Auto-Encoder RepresentationsabstractClustering is essential to many tasks in pattern recognition and computer vision. With the advent of deep learning, there is an increasing interest in learning deep unsupervised representations for clustering analysis. Many works on this domain rely on variants of auto-encoders and use the encoder outputs as representations/features for clustering. In this paper, we show that an l2normalization constraint on these representations during auto-encoder training, makes the representations more separable and compact in the Euclidean space after training. This greatly improves the clustering accuracy when k-means clustering is employed on the representations. We also propose a clustering based unsupervised anomaly detection method using l2normalized deep auto-encoder representations. We show the effect of l2normalization on anomaly detection accuracy. We further show that the proposed anomaly detection method greatly improves accuracy compared to previously proposed deep methods such as reconstruction error based anomaly detection. Çaglar Aytekin, Xingyang Ni, Francesco Cricri, Emre Aksu |
IJCNN | 3 |
| 2018 | Memory-Efficient Deep Salient Object Segmentation Networks on Gridized SuperpixelsabstractComputer vision algorithms with pixel-wise labeling tasks, such as semantic segmentation and salient object detection, have gone through a significant accuracy increase with the incorporation of deep learning. Deep segmentation methods slightly modify and fine-tune pre-trained networks that have hundreds of millions of parameters. In this work, we question the need of having such memory demanding networks for a reasonable performance in salient object segmentation. To this end, we propose a way to learn a memory-efficient network from scratch by training it only on salient object detection datasets. Our method encodes images to gridized superpixels that preserve both the object boundaries and the connectivity rules of regular pixels. This representation allows us to use convolutional neural networks that operate on regular grids. By using these encoded images, we train a memory-efficient network using only 0.048% of the number of parameters that a majority of other deep salient object detection networks have. Our method shows comparable accuracy with the state-of-the-art deep salient object detection methods and provides a much more memory-efficient alternative to them. Due to its easy deployment and small size, such a network is preferable for applications in memory limited IoT devices. Çaglar Aytekin, Xingyang Ni, Francesco Cricri, Lixin Fan, Emre Aksu |
MMSP | 3 |
| 2014 | Salient Event Detection in Basketball Mobile VideosabstractModern smartphones have become the most popular means for recording videos. In fact, thanks to their portability, smartphones allow for recording anything and at any moment of our everyday life. One common occasion is represented by sport happenings, where people often record their favourite team or players. Automatic analysis of such videos is important for enabling applications such as automatic organization, browsing and summarization of the content. This paper proposes novel algorithms for the detection of salient events in videos recorded at basketball games. The novel approach consists of jointly analyzing visual data and magnetometer data. The magnetometer data provides information about the horizontal orientation of the camera. The proposed joint analysis allows for a reduced number of false positives and for a reduced computational complexity. The algorithms are tested on data captured during real basketball games. The experimental results clearly show the advantages of the proposed approach. Francesco Cricri, Sujeet Mate, Igor D. D. Curcio, Moncef Gabbouj |
ISM | 1 |
| 2014 | Multimodal extraction of events and of information about the recording activity in user generated videosabstractIn this work we propose methods that exploit context sensor data modalities for the task of detecting interesting events and extracting high-level contextual information about the recording activity in user generated videos. Indeed, most camera-enabled electronic devices contain various auxiliary sensors such as accelerometers, compasses, GPS receivers, etc. Data captured by these sensors during the media acquisition have already been used to limit camera degradations such as shake and also to provide some basic tagging information such as the location. However, exploiting the sensor-recordings modality for subsequent higher-level information extraction such as interesting events has been a subject of rather limited research, further constrained to specialized acquisition setups. In this work, we show how these sensor modalities allow inferring information (camera movements, content degradations) about each individual video recording. In addition, we consider a multi-camera scenario, where multiple user generated recordings of a common scene (e.g., music concerts) are available. For this kind of scenarios we jointly analyze these multiple video recordings and their associated sensor modalities in order to extract higher-level semantics of the recorded media: based on the orientation of cameras we identify the region of interest of the recorded scene, by exploiting correlation in the motion of different cameras we detect generic interesting events and estimate their relative position. Furthermore, by analyzing also the audio content captured by multiple users we detect more specific interesting events. We show that the proposed multimodal analysis methods perform well on various recordings obtained in real live music performances. Francesco Cricri, Kostadin Dabov, Igor D. D. Curcio, Sujeet Mate, Moncef Gabbouj |
Multim. Tools Appl. | 1 |
| 2014 | Sport Type Classification of Mobile VideosabstractThe recent proliferation of mobile video content has emphasized the need for applications such as automatic organization and automatic editing of videos. These applications could greatly benefit from domain knowledge about the content. However, extracting semantic information from mobile videos is a challenging task, due to their unconstrained nature. We extract domain knowledge about sport events recorded by multiple users, by classifying the sport type into soccer, American football, basketball, tennis, ice-hockey, or volleyball. We adopt a multi-user and multimodal approach, where each user simultaneously captures audio-visual content and auxiliary sensor data (from magnetometers and accelerometers). Firstly, each modality is separately analyzed; then, analysis results are fused for obtaining the sport type. The auxiliary sensor data is used for extracting more discriminative spatio-temporal visual features and efficient camera motion features. The contribution of each modality to the fusion process is adapted according to the quality of the input data. We performed extensive experiments on data collected at public sport events, showing the merits of using different combinations of modalities and fusion methods. The results indicate that analyzing multimodal and multi-user data, coupled with adaptive fusion, improves classification accuracies in most tested cases, up to 95.45%. Francesco Cricri, Mikko Roininen, Jussi Leppänen, Sujeet Mate, Igor D. D. Curcio, Stefan Uhlmann, Moncef Gabbouj |
IEEE Trans. Multim. | 1 |
| 2013 | Multi-sensor fusion for sport genre classification of user generated mobile videosabstractWe present a robust multimodal approach for classifying the sport genre in videos recorded by mobile phone users at a sport event. In addition to traditional audio-visual content analysis tools, we propose to analyze auxiliary sensor data (electronic compass data and accelerometer data) captured simultaneously with the video recording. By means of machine learning techniques, we build models of visual appearance, camera motion (from auxiliary sensor data) and audio scene, which are used for classifying the data from each modality. The sport genre is obtained by fusing the information provided by the models. We propose to use the quality of each modality as an indication of its reliability. Extensive experiments were performed on real test data collected at public sport events. We provide comparisons on the use of different modality sets and fusion methods. Finally, we show how the proposed methods achieve robust classification even in the considered unconstrained scenarios. Francesco Cricri, Mikko Roininen, Sujeet Mate, Jussi Leppänen, Igor D. D. Curcio, Moncef Gabbouj |
ICME | 1 |
| 2012 | Sensor-Based Analysis of User Generated Video for Multi-camera Video Remixing
Francesco Cricri, Igor D. D. Curcio, Sujeet Mate, Kostadin Dabov, Moncef Gabbouj |
MMM | 1 |
| 2011 | We want more: human-computer collaboration in mobile social video remixing of music concertsabstractRecording and publishing mobile video clips from music concerts is popular. There is a high potential to increase the concert's perceived value when producing video remixes from individual video clips and using them socially. A digital production of a video remix is an interactive process between human and computer. However, it is not clear what the collaboration implications between human and computer are. Sami Vihavainen, Sujeet Mate, Lassi Seppälä, Francesco Cricri, Igor D. D. Curcio |
CHI | 4 |
| 2011 | Multimodal Event Detection in User Generated VideosabstractNowadays most camera-enabled electronic devices contain various auxiliary sensors such as accelerometers, gyroscopes, compasses, GPS receivers, etc. These sensors are often used during the media acquisition to limit camera degradations such as shake and also to provide some basic tagging information such as the location used in geo-tagging. Surprisingly, exploiting the sensor-recordings modality for high-level event detection has been a subject of rather limited research, further constrained to highly specialized acquisition setups. In this work, we show how these sensor modalities, alone or in combination with content-based analysis, allow inferring information about the video content. In addition, we consider a multi-camera scenario, where multiple user generated recordings of a common scene (e.g., music concerts, public events) are available. In order to understand some higher-level semantics of the recorded media, we jointly analyze the individual video recordings and sensor measurements of the multiple users. The detected semantics include generic interesting events and some more specific events. The detection exploits correlations in the camera motion and in the audio content of multiple users. We show that the proposed multimodal analysis methods perform well on various recordings obtained in real live music performances. Francesco Cricri, Kostadin Dabov, Igor D. D. Curcio, Sujeet Mate, Moncef Gabbouj |
ISM | 1 |
| 2009 | Mobile and Interactive Social Television - A virtual TV roomabstractSmart phones are becoming more and more powerful. Services that were traditionally designed for a static environment can now be implemented into mobile devices. One such service is Interactive Social TV, which allows geographically dispersed people to meet in a virtual shared space and watch TV while being able to interact with each other. This paper presents two novel architectures of a Mobile and Interactive Social TV system. In both of the architectures, the interaction is represented by a rich audio-visual media, allowing users to hear and see each other. In the first architecture, the mixing of the TV content with the interaction media is performed at the server side. In the second architecture, the mixing is performed in each client device. The issues of decoding and rendering simultaneous content and interaction media streams on a mobile device are discussed, and the related implementation is presented. Francesco Cricri, Sujeet Mate, Igor D. D. Curcio, Moncef Gabbouj |
WOWMOM | 1 |