VLDB 2026 Research / reviewers in the wild / expert
Jukka I. Ahonen
dblp:311/4165
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Evaluating the Emerging MPEG Video Coding for Machines in Semantic SegmentationabstractEmerging MPEG Video Coding for Machines (MPEG VCM) standardization activities address the growing demand for machine-to-machine visual applications, including video surveillance, autonomous driving, etc. This paper proposes an evaluation methodology tailored to MPEG VCM, with an emphasis on semantic segmentation tasks using the Pandaset dataset. This is a challenging target as standardization works must follow several limitations, such as dataset's licensing, fixed tools and software, and compliance with existing common test conditions (CTC) for the standard's development. The proposed evaluation methodology includes a step to align Pandaset and COCO labels. Two task networks, Detectron2 and Mask2Former, are used to evaluate the Rate-Performance behavior for semantic segmentation under various coding configurations. The performance of MPEG VCM is benchmarked against traditional codecs (VVC and HEVC), with a detailed analysis of MPEG VCM's coding tools. The extensive evaluations reveal interesting observations. (1) Although VCM achieved reasonably good segmentation performance, some of its developed tools, such as temporal resampling and region-of-interest coding, were not well suited for segmentation task. (2) The Hybrid NNVVC Inner Codec outperformed the VVC Inner Codec. (3) VCM's performance varies significantly for segmented classes. (4) Despite significant differences in human vision performance, VVC and HEVC exhibit relatively similar performance in machine vision. The main contributions of this work are to (1) enable evaluation of MPEG VCM in a real-world semantic segmentation use case, which is one of VCM's targeted tasks, and (2) to provide a detailed assessment of VCM's performance in semantic segmentation. Khoa Dang Pham, Farhad Pakdaman, Honglei Zhang 0001, Hamed Rezazadegan Tavakoli, Nam Le 0003, Jukka I. Ahonen, Moncef Gabbouj |
ISM | 6 |
| 2025 | Learned Image Codec with Progressive Multi-Scale Probability Model for Streaming in Unreliable Communication ChannelsabstractVideo streaming over the Internet is one of the most important applications of video compression technologies. During streaming, network congestion and packet loss can catastrophically corrupt conventional block-based codecs, yielding fully corrupted frames or partially visible content that significantly degrades user experience. Unlike block-based conventional codecs, end-to-end learned codecs operate on holistic feature representations, unlocking a new paradigm for error resilience. In this work, we leverage a progressive multi-scale entropy model to partition the bitstream into ordered data units, ensuring that any received prefix unit yields a low-quality but full-frame reconstruction. To handle missing information in the latent tensor, we introduce lightweight adapters that predict absent features before image synthesis. Experiments on JVET CTC classes show that our methods improve PSNR by up to 2.27 dB over maximum-likelihood tensor filling, at no extra bitrate cost. Honglei Zhang 0001, A. Burakhan Koyuncu, Jukka I. Ahonen, Nannan Zou, Francesco Cricri |
MMSP | 3 |
| 2025 | A Hybrid Framework Integrating End-to-End Learned Image Codec with Conventional Codec
Nannan Zou, Antti Hallapuro, Francesco Cricri, Honglei Zhang 0001, A. Burakhan Koyuncu, Jukka I. Ahonen, Miska M. Hannuksela, Esa Rahtu |
PCS | 6 |
| 2024 | Competitive Learning For Achieving Content-Specific Filters In Video Coding For MachinesabstractThis paper investigates the efficacy of jointly optimizing content-specific post-processing filters to adapt a human-oriented video/image codec into a codec suitable for machine vision tasks. By observing that artifacts produced by video/image codecs are content-dependent, we propose a novel training strategy based on competitive learning principles. This strategy assigns training samples to filters dynamically, in a fuzzy manner, which further optimizes the winning filter on the given sample. Inspired by simulated annealing optimization techniques, we employ a softmax function with a temperature variable as the weight allocation function to mitigate the effects of random initialization. Our evaluation, conducted on a system utilizing multiple post-processing filters within a Versatile Video Coding (VVC) codec framework, demonstrates the superiority of content-specific filters trained with our proposed strategies, specifically, when images are processed in blocks. Using VVC reference software VTM 12.0 as the anchor, experiments on the OpenImages dataset show an improvement in the BD-rate reduction from -41.3% and -44.6% to -42.3% and -44.7% for object detection and instance segmentation tasks, respectively, compared to independently trained filters. The statistics of the filter usage align with our hypothesis and underscore the importance of jointly optimizing filters for both content and reconstruction quality. Our findings pave the way for further improving the performance of video/image codecs. Honglei Zhang 0001, Jukka I. Ahonen, Nam Le 0003, Ruiying Yang, Francesco Cricri |
ICIP | 2 |
| 2023 | NN-VVC: Versatile Video Coding boosted by self-supervisedly learned image coding for machinesabstractThe recent progress in artificial intelligence has led to an ever-increasing usage of images and videos by machine analysis algorithms, mainly neural networks. Nonetheless, compression, storage and transmission of media have traditionally been designed considering human beings as the viewers of the content. Recent research on image and video coding for machine analysis has progressed mainly in two almost orthogonal directions. The first is represented by end-to-end (E2E) learned codecs which, while offering high performance on image coding, are not yet on par with state-of-the-art conventional video codecs and lack interoperability. The second direction considers using the Versatile Video Coding (VVC) standard or any other conventional video codec (CVC) together with pre- and post-processing operations targeting machine analysis. While the CVC-based methods benefit from interoperability and broad hardware and software support, the machine task performance is often lower than the desired level, particularly in low bitrates. This paper proposes a hybrid codec for machines called NN-VVC, which combines the advantages of an E2E-learned image codec and a CVC to achieve high performance in both image and video coding for machines. Our experiments show that the proposed system achieved up to -43.20% and -26.8% Bjøntegaard Delta rate reduction over VVC for image and video data, respectively, when evaluated on multiple different datasets and machine vision tasks. To the best of our knowledge, this is the first research paper showing a hybrid video codec that outperforms VVC on multiple datasets and multiple machine vision tasks. Jukka I. Ahonen, Nam Le 0003, Honglei Zhang 0001, Antti Hallapuro, Francesco Cricri, Hamed Rezazadegan Tavakoli, Miska M. Hannuksela, Esa Rahtu |
ISM | 1 |
| 2023 | Region of Interest Enabled Learned Image Coding for MachinesabstractImage and video coding for machines has been recently gaining more and more interest from both the industry and the research community. One successful approach is based on end-to-end (E2E) learned compression and has shown significant gains over the state-of-the-art conventional image coding methods. However, one of the remaining challenges for such E2E-learned image codecs for machines is to adaptively allocate the bits over different regions of the image, while retaining the machine vision performance. In this paper, we propose a method that leverages Regions-Of-Interest (ROIs) for bitrate allocation within a Learned Image Codec (LIC) for machines. In particular, the proposed method reduces the bits allocated for the background regions of the image by reducing the variance of the elements corresponding to the background regions in the latent representation. This results in more heavily quantized background areas, while keeping the quality of the ROI areas suitable for machine tasks. The proposed method achieves significant gains, -15.80% and -22.43% Pareto BD-rate reduction, over the baseline LIC on object detection and instance segmentation tasks, respectively. To the best of our knowledge, this is the first research paper proposing an ROI-based inference-time technology for Learned Image Coding for machines. Jukka I. Ahonen, Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Esa Rahtu |
MMSP | 1 |
| 2021 | Learned Enhancement Filters for Image Coding for MachinesabstractMachine-To-Machine (M2M) communication applications and use cases, such as object detection and instance segmentation, are becoming mainstream nowadays. As a consequence, majority of multimedia content is likely to be consumed by machines in the coming years. This opens up new challenges on efficient compression of this type of data. Two main directions are being explored in the literature, one being based on existing traditional codecs, such as the Versatile Video Coding (VVC) standard, that are optimized for human-targeted use cases, and another based on end-to-end trained neural networks. However, traditional codecs have significant benefits in terms of interoperability, real-time decoding, and availability of hardware implementations over end-to-end learned codecs. Therefore, in this paper, we propose learned post-processing filters that are targeted for enhancing the performance of machine vision tasks for images reconstructed by the VVC codec. The proposed enhancement filters provide significant improvements on the target tasks compared to VVC coded images. The conducted experiments show that the proposed post-processing filters provide about 45% and 49% Bjøntegaard Delta Rate gains over VVC in instance segmentation and object detection tasks, respectively. Jukka I. Ahonen, Ramin Ghaznavi Youvalari, Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Miska M. Hannuksela, Esa Rahtu |
ISM | 1 |