Honglei Zhang 0001

dblp:40/2812-1 · DBLP profile ↗
← Back
32ranked-venue papers
7as first author
28since 2021 · last 2025
0000-0002-8229-852XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32 · 7 first-author · 28 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Evaluating the Emerging MPEG Video Coding for Machines in Semantic Segmentation
abstract
Emerging MPEG Video Coding for Machines (MPEG VCM) standardization activities address the growing demand for machine-to-machine visual applications, including video surveillance, autonomous driving, etc. This paper proposes an evaluation methodology tailored to MPEG VCM, with an emphasis on semantic segmentation tasks using the Pandaset dataset. This is a challenging target as standardization works must follow several limitations, such as dataset's licensing, fixed tools and software, and compliance with existing common test conditions (CTC) for the standard's development. The proposed evaluation methodology includes a step to align Pandaset and COCO labels. Two task networks, Detectron2 and Mask2Former, are used to evaluate the Rate-Performance behavior for semantic segmentation under various coding configurations. The performance of MPEG VCM is benchmarked against traditional codecs (VVC and HEVC), with a detailed analysis of MPEG VCM's coding tools. The extensive evaluations reveal interesting observations. (1) Although VCM achieved reasonably good segmentation performance, some of its developed tools, such as temporal resampling and region-of-interest coding, were not well suited for segmentation task. (2) The Hybrid NNVVC Inner Codec outperformed the VVC Inner Codec. (3) VCM's performance varies significantly for segmented classes. (4) Despite significant differences in human vision performance, VVC and HEVC exhibit relatively similar performance in machine vision. The main contributions of this work are to (1) enable evaluation of MPEG VCM in a real-world semantic segmentation use case, which is one of VCM's targeted tasks, and (2) to provide a detailed assessment of VCM's performance in semantic segmentation.
Khoa Dang Pham, Farhad Pakdaman, Honglei Zhang 0001, Hamed Rezazadegan Tavakoli, Nam Le 0003, Jukka I. Ahonen, Moncef Gabbouj
ISM3
2025 Learned Image Codec with Progressive Multi-Scale Probability Model for Streaming in Unreliable Communication Channels
abstract
Video streaming over the Internet is one of the most important applications of video compression technologies. During streaming, network congestion and packet loss can catastrophically corrupt conventional block-based codecs, yielding fully corrupted frames or partially visible content that significantly degrades user experience. Unlike block-based conventional codecs, end-to-end learned codecs operate on holistic feature representations, unlocking a new paradigm for error resilience. In this work, we leverage a progressive multi-scale entropy model to partition the bitstream into ordered data units, ensuring that any received prefix unit yields a low-quality but full-frame reconstruction. To handle missing information in the latent tensor, we introduce lightweight adapters that predict absent features before image synthesis. Experiments on JVET CTC classes show that our methods improve PSNR by up to 2.27 dB over maximum-likelihood tensor filling, at no extra bitrate cost.
Honglei Zhang 0001, A. Burakhan Koyuncu, Jukka I. Ahonen, Nannan Zou, Francesco Cricri
MMSP1
2025 Task Enhancement Tiles for Ultra Lightweight Post-processing in Visual Coding for Machines
abstract
The proliferation of automated visual analysis calls for compression methods tailored to the unique requirements of Video Coding for Machines (VCM). In this paper, we propose a computationally lightweight post-processing method that is based on a learned component referred to as a task enhancement tile (TET). A TET is spatially tiled over the reconstructed visual data and added to it element-wise. It only requires one addition per pixel in each color channel before the machine task can be applied. Our results with the VVC test model (VTM) demonstrate coding gains of up to 39.0% for object detection and 29.2% for instance segmentation on image datasets, while evaluation on a video dataset shows gains of up to 35.2% for object detection, relative to the VTM anchor. The proposed solution also offers extremely low computational cost, preservation of human-viewable content, full compliance with video coding standards, no requirement for side information transmission from encoder to decoder, and generalization across tasks, models, and encoding parameters.
Tero Partanen, Alban Marie, Rudolf Kortelahti, Alexandre Mercat, Jarno Vanne, Miska M. Hannuksela, Honglei Zhang 0001, Alireza Aminlou, Francesco Cricri
PCS7
2025 A Hybrid Framework Integrating End-to-End Learned Image Codec with Conventional Codec
Nannan Zou, Antti Hallapuro, Francesco Cricri, Honglei Zhang 0001, A. Burakhan Koyuncu, Jukka I. Ahonen, Miska M. Hannuksela, Esa Rahtu
PCS4
2024 Competitive Learning For Achieving Content-Specific Filters In Video Coding For Machines
abstract
This paper investigates the efficacy of jointly optimizing content-specific post-processing filters to adapt a human-oriented video/image codec into a codec suitable for machine vision tasks. By observing that artifacts produced by video/image codecs are content-dependent, we propose a novel training strategy based on competitive learning principles. This strategy assigns training samples to filters dynamically, in a fuzzy manner, which further optimizes the winning filter on the given sample. Inspired by simulated annealing optimization techniques, we employ a softmax function with a temperature variable as the weight allocation function to mitigate the effects of random initialization. Our evaluation, conducted on a system utilizing multiple post-processing filters within a Versatile Video Coding (VVC) codec framework, demonstrates the superiority of content-specific filters trained with our proposed strategies, specifically, when images are processed in blocks. Using VVC reference software VTM 12.0 as the anchor, experiments on the OpenImages dataset show an improvement in the BD-rate reduction from -41.3% and -44.6% to -42.3% and -44.7% for object detection and instance segmentation tasks, respectively, compared to independently trained filters. The statistics of the filter usage align with our hypothesis and underscore the importance of jointly optimizing filters for both content and reconstruction quality. Our findings pave the way for further improving the performance of video/image codecs.
Honglei Zhang 0001, Jukka I. Ahonen, Nam Le 0003, Ruiying Yang, Francesco Cricri
ICIP1
2024 Feasibility Study of Multi-Layer VVC Coding Scheme for Hybrid Machine-Human Consumption
abstract
The proliferation of machine vision applications necessitates developing more efficient visual data compression schemes for machine consumption. However, numerous automated use cases still require keeping humans in the loop, leading to the need for a machine-optimized video streaming with the option for human supervision. This paper investigates the feasibility of using the multi-layer coding approach of the emerging Versatile Video Coding (VVC) standard to create favorable conditions for hybrid machine-human consumption. We introduce a multi-layer coding scheme, where the base layer (BL) is optimized for machines and the enhancement layer (EL) complements the stream for human vision. Our results demonstrate that the bitrate of the proposed multi-layer stream (BL + EL) is, on average, 11% higher than that of a single-layer VVC. However, the more compact BL yields overall bandwidth savings as long as the EL is required less than 80% of the time.
Jaakko Laitinen, Tero Partanen, Alexandre Mercat, Jarno Vanne, Miska M. Hannuksela, Honglei Zhang 0001, Alireza Aminlou, Francesco Cricri
ICME6
2024 Luma Range Scaling for Enhanced VVC Efficiency in Video Coding for Machines
abstract
Recent years have shown significant growth in video data traffic for machine vision applications, catalyzing new standardization efforts in video coding for machines (VCM). These activities focus on compressing images and videos for machine vision tasks, rather than for human viewing. In this work, we propose a novel method that scales down the luma range to enhance the coding efficiency of Versatile Video Coding (VVC) for machine consumption. This method results in a lower bitrate after encoding and has only minimal adverse effects on the accuracy of machine vision tasks. In our experiments, we down-scale the luma channel of the input video using luma-scaling factors from 0.2 to 0.9 and evaluate coding results with optional back-scaling to the original range before machine vision tasks. Our results with the VVC Test Model (VTM) demonstrate that the proposed technique achieves coding gain of up to 37.9%and 46.1% for the same object detection and tracking accuracy, respectively.
Tero Partanen, Alban Marie, Alexandre Mercat, Jarno Vanne, Miska M. Hannuksela, Honglei Zhang 0001, Alireza Aminlou, Francesco Cricri
MMSP6
2024 Packed Regions Information SEI Message
abstract
Specific regions of interest (ROIs) within a video are of greater interest than the remainder of the video for many use cases. The Packed Regions Information (PRI) Supplemental Enhancement Information (SEI) message enables packing of rectangular ROIs from an original picture into a smaller resolution picture for video coding, reducing pixel rate and bitrate. The SEI message signals metadata describing the size and position of the ROIs in the coded picture and in the original picture. Decoders may use the metadata to reconstruct target pictures at the original resolution from the decoded pictures containing the packed regions. The PRI SEI message is under consideration for potential inclusion in a future version of the Versatile Supplemental Enhanced Information (VSEI) standard. Experimental results are provided for use of the PRI SEI message with test conditions for machine analysis of coded video content, showing reductions in bitrate and pixel rate.
Jill M. Boyce, Miska M. Hannuksela, Honglei Zhang 0001, Antti Hallapuro
VCIP3
2023 NN-VVC: Versatile Video Coding boosted by self-supervisedly learned image coding for machines
abstract
The recent progress in artificial intelligence has led to an ever-increasing usage of images and videos by machine analysis algorithms, mainly neural networks. Nonetheless, compression, storage and transmission of media have traditionally been designed considering human beings as the viewers of the content. Recent research on image and video coding for machine analysis has progressed mainly in two almost orthogonal directions. The first is represented by end-to-end (E2E) learned codecs which, while offering high performance on image coding, are not yet on par with state-of-the-art conventional video codecs and lack interoperability. The second direction considers using the Versatile Video Coding (VVC) standard or any other conventional video codec (CVC) together with pre- and post-processing operations targeting machine analysis. While the CVC-based methods benefit from interoperability and broad hardware and software support, the machine task performance is often lower than the desired level, particularly in low bitrates. This paper proposes a hybrid codec for machines called NN-VVC, which combines the advantages of an E2E-learned image codec and a CVC to achieve high performance in both image and video coding for machines. Our experiments show that the proposed system achieved up to -43.20% and -26.8% Bjøntegaard Delta rate reduction over VVC for image and video data, respectively, when evaluated on multiple different datasets and machine vision tasks. To the best of our knowledge, this is the first research paper showing a hybrid video codec that outperforms VVC on multiple datasets and multiple machine vision tasks.
Jukka I. Ahonen, Nam Le 0003, Honglei Zhang 0001, Antti Hallapuro, Francesco Cricri, Hamed Rezazadegan Tavakoli, Miska M. Hannuksela, Esa Rahtu
ISM3
2023 Region of Interest Enabled Learned Image Coding for Machines
abstract
Image and video coding for machines has been recently gaining more and more interest from both the industry and the research community. One successful approach is based on end-to-end (E2E) learned compression and has shown significant gains over the state-of-the-art conventional image coding methods. However, one of the remaining challenges for such E2E-learned image codecs for machines is to adaptively allocate the bits over different regions of the image, while retaining the machine vision performance. In this paper, we propose a method that leverages Regions-Of-Interest (ROIs) for bitrate allocation within a Learned Image Codec (LIC) for machines. In particular, the proposed method reduces the bits allocated for the background regions of the image by reducing the variance of the elements corresponding to the background regions in the latent representation. This results in more heavily quantized background areas, while keeping the quality of the ROI areas suitable for machine tasks. The proposed method achieves significant gains, -15.80% and -22.43% Pareto BD-rate reduction, over the baseline LIC on object detection and instance segmentation tasks, respectively. To the best of our knowledge, this is the first research paper proposing an ROI-based inference-time technology for Learned Image Coding for machines.
Jukka I. Ahonen, Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Esa Rahtu
MMSP3
2023 Overfitting NN loop-filters in video coding
abstract
Overfitting is usually regarded as a negative condition since it impairs the generalisation power of a model. Nevertheless, overfitting a Neural Network (NN) on test data may be advantageous to improve the compression efficiency of image/video coding tools and systems. Previous research has demonstrated the benefits of NN overfitting for post-processing operations, i.e. post-filters, but not yet for actual decoding tools. Generally, the NN is overfitted on test data at the encoder end, and the weight update is coded and sent to the decoder end along the image/video bitstream. The proposed approach follows this strategy. In particular, the overfitting of the Low Operation Point (LOP) loop-filter in NN-based Video Coding (NNVC) software is studied. The overall approach yields Bjøntegaard Delta rate (BD-rate) of -7.74%, -13.73% and -12.49%, for the Y, U and V components, respectively. Out of these coding gains, 1.21%, 6.43% and 5.52%, for the Y, U and V components, are attributed to the overfitting. The boost in the coding gains comes with only 1.5% more complexity, due to the multiplier parameters introduced during the overfitting.
Ruiying Yang, María Santamaría 0001, Francesco Cricri, Honglei Zhang 0001, Jani Lainema, Ramin Ghaznavi Youvalari, Miska M. Hannuksela, Tapio Elomaa
VCIP4
2022 Content-Adaptive Neural Network Post-Processing Filter with NNR-Coded Weight-Updates
abstract
Neural Network (NN) filters improve the perceptual quality of reconstructed videos by reducing compression artefacts. For content adaptation, a few NN-filters use over-fitting. As the adaptation signal is a weight-update, compression is required to minimise significant bitrate overheads. Most approaches, however, use generic data compression algorithms, which are inadequate for coding NN weight-updates. This work introduces a content-adaptive NN post-processing filter with weight-updates coded using the Neural Network compression and Representation (NNR) standard. The bitrate overhead is further decreased by over-fitting only a subset of weights, selected via energy-based analysis. The proposed filter saved about 4.57% (Y), 10.33% (Cb), 6.53% (Cr) Bjøntegaard Delta rate (BD-rate) on top of the Versatile Video Coding (VVC) Test Model (VTM) 11.0 with NN-based Video Coding (NNVC) 1.0, in Random Access (RA) configuration. Compared to the non-over-fitted NN, the performance was doubled; and compared to 7z, NNR reduced the bitrate of the weight-update by ∼64%.
María Santamaría 0001, Francesco Cricri, Jani Lainema, Ramin Ghaznavi Youvalari, Honglei Zhang 0001, Miska M. Hannuksela
ICIP5
2022 Bridging the Gap Between Image Coding for Machines and Humans
abstract
Image coding for machines (ICM) aims at reducing the bitrate required to represent an image while minimizing the drop in machine vision analysis accuracy. In many use cases, such as surveillance, it is also important that the visual quality is not drastically deteriorated by the compression process. Recent works on using neural network (NN) based ICM codecs have shown significant coding gains against traditional methods; however, the decompressed images, especially at low bitrates, often contain checkerboard artifacts. We propose an effective decoder finetuning scheme based on adversarial training to significantly enhance the visual quality of ICM codecs, while preserving the machine analysis accuracy, without adding extra bitcost or parameters at the inference phase. The results show complete removal of the checkerboard artifacts at the negligible cost of −1.6% relative change in task performance score. In the cases where some amount of artifacts is tolerable, such as when machine consumption is the primary target, this technique can enhance both pixel-fidelity and feature-fidelity scores without losing task performance.
Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Hamed Rezazadegan Tavakoli, Emre Aksu, Miska M. Hannuksela, Esa Rahtu
ICIP2
2022 Stochastic Binary-Ternary Quantization for Communication Efficient Federated Computation
abstract
A stochastic binary-ternary (SBT) quantization approach is introduced for communication efficient federated computation; form of collaborative computing where locally trained models are exchanged between institutes. Communication of deep neural network models could be highly inefficient due to their large size. This motivates model compression in which quantization is an important step. Two well-known quantization algorithms are binary and ternary quantization. The first leads into good compression, sacrificing accuracy. The second provides good accuracy with less compression. To better benefit from trade-off between accuracy and compression, we propose an algorithm to stochastically switch between binary and ternary quantization. By combining with uniform quantization, we further extend the proposed algorithm to a hierarchical method which results in even better compression without sacrificing the accuracy. We tested the proposed algorithm using Neural network Compression Test Model (NCTM) provided by MPEG community. Our results demonstrate that the hierarchical variant of the proposed algorithm outperforms other quantization algorithms in term of compression, while maintaining the accuracy competitive to that provided by other methods.
Rangu Goutham, Homayun Afrabandpey, Francesco Cricri, Honglei Zhang 0001, Emre Aksu, Miska M. Hannuksela, Hamed Rezazadegan Tavakoli
ICIP4
2022 Adaptive Multi-Scale Progressive Probability Model for Lossless Image Compression
abstract
Domain adaptation is an efficient technique to improve the performance of a system by adapting a pre-trained model to the given input data. The adaptation technique has been generally applied in conventional video codecs. For neural network-based systems, the encoder may adapt the decoder to the input data by fine-tuning a pre-trained model present at the decoder side. The weight update is then transferred to the decoder and the updated model is used to decode the bitstream. However, due to the large number of parameters in deep neural networks, the overhead of the weight update may diminish the gain from the adaptation technique. In recent years, various methods have been proposed to reduce the overhead without significantly compromising the gain. In this paper, we propose an adaptive multi-scale progressive probability model for lossless image compression. The proposed method uses the data that has already been processed at the inference stage to fine-tune the probability model. Importantly, the decoder can apply the fine-tuning by itself resulting a small adaptation overhead to help the decoder in performing the fine-tuning. The proposed method achieves up to 0.28 bits-per-pixel (BPP) reduction on four benchmark datasets compared to the state-of-the-art method.
Honglei Zhang 0001, Francesco Cricri, Nannan Zou, Hamed Rezazadegan Tavakoli, Miska M. Hannuksela
ICIP1
2022 TMD: Transformed Mesh Decoder for Mesh Animation
abstract
Easy and fast animation of 3D characters is attractive for both gaming and entertainment applications. 3D mesh must be properly rigged and skinned in order to create seamless animation. This process can be time consuming and requires deep knowledge of appropriate software based on kinematic animation. In this work, we present a fast and lightweight deep neural model to automate 3D human animation using skeletal representation from 2D image pose, i.e., joints in 2D space. We accomplish this using Transformed Mesh Decoder (TMD), which is a novel layer for convolutional neural networks. To train the network, we generate a large and diverse dataset using Skinned Multi-Person Linear (SMPL) model. Experiment shows that our method is effective when compared to both the ground truth and state-of-the art linear blend skinning that require manually painted skinning weights for accurate result. The animation process is fast and can achieve approximately 10-15fps in practice. The proposed method is simple which opens the possibility for future improvement in real-time application.
Peter Fasogbon, Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Emre Aksu
ICPR2
2022 Low-precision post-filtering in video coding
abstract
Neural Networks (NNs) have demonstrated their effectiveness in tackling challenges involving multimedia content. In the video coding field, NNs are actively exploited as novel tools that complement conventional signal processing tools, as well as end-to-end coding solutions. Since NNs use commonly floating-point arithmetic, different results may be generated in different computing environments, leading to discrepancies and even corrupted reconstructions. Accordingly, this issue is solved by employing fixed-point arithmetic instead. This paper studies the quantisation of a 32-bit floating-point (float32) NN post-filter to 32-bit fixed-point (int32) and 16-bit fixed-point (int16). On top of the VVC Test Model (VTM) 11.0 NN-based Video Coding (NNVC) 1.0, the coding gains of the float32 post-filter are 5.01% (Y), 18.95% (Cb) and 17.33% Cr. Compared to the float32 inference, the quantised models produce coding losses: 0.01% (Y, Cb and Cr) for the int32 inference and 0.45% (Y), 1.97% (Cb) and 1.19% (Cr) for the int16 inference. Nevertheless, the fixed-point approaches achieve bit exact matches in different computing environments. Moreover, the decoding time with int16 is about half the decoding time of int32.
Ruiying Yang, María Santamaría 0001, Francesco Cricri, Honglei Zhang 0001, Jani Lainema, Ramin Ghaznavi Youvalari, Miska M. Hannuksela
ISM4
2022 The Lottery Ticket Adaptation for Neural Video Coding
abstract
Recently, learning based video compression methods have attracted increasing attention. However, most learning based video codecs are not adaptive to different video contents. Though adaptation at inference time is a solution to tackle this issue, adapting all the codec’s parameters is computationally expensive and brings heavy bitrate overhead. The recently proposed Lottery Ticket Hypothesis (LTH) states that an over-parameterized neural network contains smaller subnetworks (winning tickets) that can match the performance of the original network. In this paper, we present a novel lottery-ticket adaptation technique on decoder-side multiplicative parameters of a neural network, transferring the concept of winning lottery tickets to video compression tasks. At inference time, the winning multiplicative parameters are overfitted, compressed, and signaled together with encoded frames for decoding. We show that our approach outperforms the Versatile Video Coding (VVC) standard in the Multiscale Structural Similarity (MS-SSIM) at a low bitrate on both the UVG and JVET sequences. To the best of our knowledge, this is the first attempt to apply LTH in the video compression domain. Also, this is the first published end-to-end learned video codec working directly on YUV format, which outperforms VVC on UVG and JVET datasets in MS-SSIM.
Nannan Zou, Francesco Cricri, Honglei Zhang 0001, Hamed Rezazadegan Tavakoli, Miska M. Hannuksela, Esa Rahtu
ISM3
2022 On the Importance of Temporal Dependencies of Weight Updates in Communication Efficient Federated Learning
abstract
This paper studies the effect of exploiting temporal dependency of successive weight updates on compressing communications in Federated Learning (FL). For this, we propose residual coding for FL, which utilizes temporal dependencies by communicating compressed residuals of the weight updates whenever they are beneficial to bandwidth. We further consider Temporal Context Adaptation (TCA) which compares co-located elements of consecutive weight updates to select optimal setting for compression of bitstream in DeepCABAC encoder. Following experimental settings of MPEG standard on Neural Network Compression (NNC), we demonstrate that both temporal dependency based technologies reduce communication overhead, where the maximum reduction is obtained using both technologies, simultaneously.
Homayun Afrabandpey, Rangu Goutham, Honglei Zhang 0001, Francesco Criri, Emre Aksu, Hamed Rezazadegan Tavakoli
VCIP3
2021 Image Coding For Machines: an End-To-End Learned Approach
abstract
Over recent years, deep learning-based computer vision systems have been applied to images at an ever-increasing pace, oftentimes representing the only type of consumption for those images. Given the dramatic explosion in the number of images generated per day, a question arises: how much better would an image codec targeting machine-consumption perform against state-of-the-art codecs targeting human-consumption? In this paper, we propose an image codec for machines which is neural network (NN) based and end-to-end learned. In particular, we propose a set of training strategies that address the delicate problem of balancing competing loss functions, such as computer vision task losses, image distortion losses, and rate loss. Our experimental results show that our NN-based codec outperforms the state-of-the-art Versa-tile Video Coding (VVC) standard on the object detection and instance segmentation tasks, achieving -37.87% and -32.90% of BD-rate gain, respectively, while being fast thanks to its compact size. To the best of our knowledge, this is the first end-to-end learned machine-targeted image codec.
Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Esa Rahtu
ICASSP2
2021 Mind The Structure: Adopting Structural Information For Deep Neural Network Compression
abstract
Deep neural networks have huge number of parameters and require large number of bits for representation. This hinders their adoption in decentralized environments where model transfer among different parties is a characteristic of the environment while the communication bandwidth is limited. Parameter quantization is a compression approach to address this challenge by reducing the number of bits required to represent a model, e.g. a neural network. However, majority of existing neural network quantization methods do not exploit structural information of layers and parameters during quantization. In this paper, focusing on Convolutional Neural Networks (CNNs), we present a novel quantization approach by employing the structural information of neural network layers and their corresponding parameters. Starting from a pre-trained CNN, we categorize network parameters into different groups based on the similarity of their layers and their spatial structure. Parameters of each group are independently clustered and the centroid of each cluster is used as representative for all parameters in the cluster. Finally, the centroids and the cluster indexes of the parameters are used as a compact representation of the parameters. Experiments with two different tasks, i.e., acoustic scene classification and image compression, demonstrate the effectiveness of the proposed approach.
Homayun Afrabandpey, Anton Muravev, Hamed Rezazadegan Tavakoli, Honglei Zhang 0001, Francesco Cricri, Moncef Gabbouj, Emre Aksu
ICIP4
2021 Hybrid Pruning And Sparsification
abstract
A hybrid approach based on the combination of saliency-based neural pruning and regularization-based sparsification is proposed. We propose using a graph diffusion process for determining the neuron importance for pruning. Then, we use a regularization loss based on weighted $L_{1}-$norm and $L_{2}-$norm during fine-tuning to recover the lost performance. This is followed by a threshold step to further impose sparsification. We demonstrate such a hybrid approach achieves significantly better performance in comparison to purely regularization-based sparsification for large neural networks. To this end, we assessed our proposed method on three tasks, including: image classification (3 network architectures), audio classification and image compression.
Hamed Rezazadegan Tavakoli, Joachim Wabnig, Francesco Cricri, Honglei Zhang 0001, Emre Aksu, Iraj Saniee
ICIP4
2021 Learned Image Coding for Machines: A Content-Adaptive Approach
abstract
Today, according to the Cisco Annual Internet Report (2018-2023), the fastest-growing category of Internet traffic is machine-to-machine communication. In particular, machine-to-machine communication of images and videos represents a new challenge and opens up new perspectives in the context of data compression. One possible solution approach consists of adapting current human-targeted image and video coding standards to the use case of machine consumption. Another approach consists of developing completely new compression paradigms and architectures for machine-to-machine communications. In this paper, we focus on image compression and present an inference-time content-adaptive fine-tuning scheme that optimizes the latent representation of an end-to-end learned image codec, aimed at improving the compression efficiency for machine-consumption. The conducted experiments targeting instance segmentation task network show that our online finetuning brings an average bitrate saving (BD-rate) of -3.66% with respect to our pretrained image codec. In particular, at low bitrate points, our proposed method results in a significant bitrate saving of -9.85%. Overall, our pretrained-and-then-finetuned system achieves - 30.54% BD-rate over the state-of-the-art image/video codec Versatile Video Coding (VVC) on instance segmentation.
Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Hamed Rezazadegan Tavakoli, Esa Rahtu
ICME2
2021 Learned Enhancement Filters for Image Coding for Machines
abstract
Machine-To-Machine (M2M) communication applications and use cases, such as object detection and instance segmentation, are becoming mainstream nowadays. As a consequence, majority of multimedia content is likely to be consumed by machines in the coming years. This opens up new challenges on efficient compression of this type of data. Two main directions are being explored in the literature, one being based on existing traditional codecs, such as the Versatile Video Coding (VVC) standard, that are optimized for human-targeted use cases, and another based on end-to-end trained neural networks. However, traditional codecs have significant benefits in terms of interoperability, real-time decoding, and availability of hardware implementations over end-to-end learned codecs. Therefore, in this paper, we propose learned post-processing filters that are targeted for enhancing the performance of machine vision tasks for images reconstructed by the VVC codec. The proposed enhancement filters provide significant improvements on the target tasks compared to VVC coded images. The conducted experiments show that the proposed post-processing filters provide about 45% and 49% Bjøntegaard Delta Rate gains over VVC in instance segmentation and object detection tasks, respectively.
Jukka I. Ahonen, Ramin Ghaznavi Youvalari, Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Miska M. Hannuksela, Esa Rahtu
ISM4
2021 Content-adaptive convolutional neural network post-processing filter
abstract
Neural Network (NN)-based coding techniques are being developed for hybrid video coding schemes, such as the Versatile Video Coding (VVC) standard. In-loop filters and postprocessing filters are two types of coding tools that aim to improve the visual quality of the reconstructed content. These tools are usually trained on large video or image datasets with varying content, but they are rarely adaptive to different content types. This problem is addressed with the proposed content-adaptive Convolutional Neural Network (CNN) post-processing filter. The proposed approach is content-adaptive in two ways. Firstly, a relatively simple CNN is pre-trained on a general video dataset and then fine-tuned on the video to be coded. Since only the bias terms of the CNN are fine-tuned, the signalling overhead is reduced. Secondly, a scaling factor indicates the influence of the CNN post-processing filter on the final reconstruction. The CNN post-processing filter is evaluated on top of VVC Test Model (VTM) 11.0 with NN-based Video Coding (NNVC) 1.0 and, overall, it can save 2.37% (Y), 3.63% (U), 2.24% (V) Bjøntegaard Delta rate (BD-rate) in the Random Access (RA) configuration.
María Santamaría 0001, Yat-Hong Lam, Francesco Cricri, Jani Lainema, Ramin Ghaznavi Youvalari, Honglei Zhang 0001, Miska M. Hannuksela, Esa Rahtu, Moncef Gabbouj
ISM6
2021 Enhancing Image Coding for Machines with Compressed Feature Residuals
abstract
As computer vision technologies have tremendously improved over the last decade, videos and images are often consumed by machines instead of humans which are the main target for traditional video codecs. In many use cases, although machines are the main consumers, human involvement is also required, or even mandatory. In this paper, we propose a novel image coding technique targeted for machines, while maintaining the capability for human consumption. Our proposed codec generates two bitstreams: one bitstream from a traditional codec, referred to as human bitstream, optimized for human consumption; the other bitstream, referred to as machine bitstream, generated from an end-to-end learned neural network-based codec and optimized for machine tasks. Instead of working on the image domain, the proposed machine bitstream is derived from feature residuals – the difference between the features extracted from the input image and the features extracted from the reconstructed image generated by the traditional codec. With the help of the machine bitstream, we can significantly improve machine task performance in the low bitrate range. Our system beats the state-of-the-art traditional codec, the Versatile Video Coding (VVC/H.266), achieving −40.5% in Bjontegaard delta bitrate reduction on average for bitrates up to 0.07 BPP.
Joni Seppälä, Honglei Zhang 0001, Nam Le 0003, Ramin Ghaznavi Youvalari, Francesco Cricri, Hamed Rezazadegan Tavakoli, Emre Aksu, Miska M. Hannuksela, Esa Rahtu
ISM2
2021 Adaptation and Attention for Neural Video Coding
abstract
Neural image coding represents now the state-of-the-art image compression approach. However, a lot of work is still to be done in the video domain. In this work, we propose an end-to-end learned video codec that introduces several architectural novelties as well as training novelties, revolving around the concepts of adaptation and attention. Our codec is organized as an intra-frame codec paired with an inter-frame codec. As one architectural novelty, we propose to train the inter-frame codec model to adapt the motion estimation process based on the resolution of the input video. A second architectural novelty is a new neural block that combines concepts from split-attention based neural networks and from DenseNets. Finally, we propose to overfit a set of decoder-side multiplicative parameters at inference time. Through ablation studies and comparisons to prior art, we show the benefits of our proposed techniques in terms of coding gains. We compare our codec to VVC/H.266 and RLVC, which represent the state-of-the-art traditional and end-to-end learned codecs, respectively, and to the top performing end-to-end learned approach in 2021 CLIC competition, E2E_T_OL. Our codec clearly outperforms E2E_T_OL, and compare favorably to VVC and RLVC in some settings.
Nannan Zou, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Hamed Rezazadegan Tavakoli, Jani Lainema, Emre Aksu, Miska M. Hannuksela, Esa Rahtu
ISM2
2021 Learn to overfit better: finding the important parameters for learned image compression
abstract
For most machine learning systems, overfitting is an undesired behavior. However, overfitting a model to a test image or a video at inference time is a favorable and effective technique to improve the coding efficiency of learning-based image and video codecs. At the encoding stage, one or more neural networks that are part of the codec are finetuned using the input image or video to achieve a better coding performance. The encoder en-codes the input content into a content bitstream. If the finetuned neural network is part (also) of the decoder, the encoder signals the weight update of the finetuned model to the decoder along with the content bitstream. At the decoding stage, the decoder first updates its neural network model according to the received weight update, and then proceeds with decoding the content bitstream. Since a neural network contains a large number of parameters, compressing the weight update is critical to reducing bitrate overhead. In this paper, we propose learning-based methods to find the important parameters to be overfitted, in terms of rate-distortion performance. Based on simple distribution models for variables in the weight update, we derive two objective functions. By optimizing the proposed objective functions, the importance scores of the parameters can be calculated and the important parameters can be determined. Our experiments on lossless image compression codec show that the proposed method significantly outperforms a prior-art method where overfitted parameters were selected based on heuristics. Furthermore, our technique improved the compression performance of the state-of-the-art lossless image compression codec by 0.1 bit per pixel.
Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, María Santamaría 0001, Yat-Hong Lam, Miska M. Hannuksela
VCIP1
2020 Lossless Image Compression Using a Multi-scale Progressive Statistical Model
Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Nannan Zou, Emre Aksu, Miska M. Hannuksela
ACCV (3)1
2020 L2C - Learning to Learn to Compress
abstract
In this paper we present an end-to-end meta-learned system for image compression. Traditional machine learning based approaches to image compression train one or more neural network for generalization performance. However, at inference time, the encoder or the latent tensor output by the encoder can be optimized for each test image. This optimization can be regarded as a form of adaptation or benevolent overfitting to the input content. In order to reduce the gap between training and inference conditions, we propose a new training paradigm for learned image compression, which is based on meta-learning. In a first phase, the neural networks are trained normally. In a second phase, the Model-Agnostic Meta-learning approach is adapted to the specific case of image compression, where the inner-loop performs latent tensor overfitting, and the outer loop updates both encoder and decoder neural networks based on the overfitting performance. Furthermore, after meta-learning, we propose to overfit and cluster the bias terms of the decoder on training image patches, so that at inference time the optimal content-specific bias terms can be selected at encoder-side. Finally, we propose a new probability model for lossless compression, which combines concepts from both multi-scale and super-resolution probability model approaches. We show the benefits of all our proposed ideas via carefully designed experiments.
Nannan Zou, Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Jani Lainema, Miska M. Hannuksela, Emre Aksu, Esa Rahtu
MMSP2
2018 Feature Dimensionality Reduction with Graph Embedding and Generalized Hamming Distance
abstract
Principal component analysis (PCA) and linear discriminant analysis (LDA) are the most well-known methods to reduce the dimensionality of feature vectors. However, both methods face challenges when used on multilabel data - each data point may be associated to multiple labels. PCA does not take advantage of label information thus the performance is sacrificed. LDA can exploit class information for multiclass data, but cannot be directly applied to multilabel problems. In this paper, we propose a novel dimensionality reduction method for multilabel data. We first introduce the generalized Hamming distance that measures the distance of two data points in the label space. Then the proposed distance is used in the graph embedding framework for feature dimension reduction. We verified the proposed method using three multilabel benchmark datasets and one large image dataset. The results show that the proposed feature dimensionality reduction method consistently outperforms PCA and other competing methods.
Honglei Zhang 0001, Moncef Gabbouj
ICIP1
2017 A k-nearest neighbor multilabel ranking algorithm with application to content-based image retrieval
abstract
Multilabel ranking is an important machine learning task with many applications, such as content-based image retrieval (CBIR). However, when the number of labels is large, traditional algorithms are either infeasible or show poor performance. In this paper, we propose a simple yet effective multilabel ranking algorithm that is based on k-nearest neighbor paradigm. The proposed algorithm ranks labels according to the probabilities of the label association using the neighboring samples around a query sample. Different from traditional approaches, we take only positive samples into consideration and determine the model parameters by directly optimizing ranking loss measures. We evaluated the proposed algorithm using four popular multilabel datasets. The proposed algorithm achieves equivalent or better performance than other instance-based learning algorithms. When applied to a CBIR system with a dataset of 1 million samples and over 190 thousand labels, which is much larger than any other multilabel datasets used earlier, the proposed algorithm clearly outperforms the competing algorithms.
Honglei Zhang 0001, Serkan Kiranyaz, Moncef Gabbouj
ICASSP1