Alban Marie

dblp:296/4451 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2025
0009-0002-6154-6974ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Task Enhancement Tiles for Ultra Lightweight Post-processing in Visual Coding for Machines
abstract
The proliferation of automated visual analysis calls for compression methods tailored to the unique requirements of Video Coding for Machines (VCM). In this paper, we propose a computationally lightweight post-processing method that is based on a learned component referred to as a task enhancement tile (TET). A TET is spatially tiled over the reconstructed visual data and added to it element-wise. It only requires one addition per pixel in each color channel before the machine task can be applied. Our results with the VVC test model (VTM) demonstrate coding gains of up to 39.0% for object detection and 29.2% for instance segmentation on image datasets, while evaluation on a video dataset shows gains of up to 35.2% for object detection, relative to the VTM anchor. The proposed solution also offers extremely low computational cost, preservation of human-viewable content, full compliance with video coding standards, no requirement for side information transmission from encoder to decoder, and generalization across tasks, models, and encoding parameters.
Tero Partanen, Alban Marie, Rudolf Kortelahti, Alexandre Mercat, Jarno Vanne, Miska M. Hannuksela, Honglei Zhang 0001, Alireza Aminlou, Francesco Cricri
PCS2
2024 Luma Range Scaling for Enhanced VVC Efficiency in Video Coding for Machines
abstract
Recent years have shown significant growth in video data traffic for machine vision applications, catalyzing new standardization efforts in video coding for machines (VCM). These activities focus on compressing images and videos for machine vision tasks, rather than for human viewing. In this work, we propose a novel method that scales down the luma range to enhance the coding efficiency of Versatile Video Coding (VVC) for machine consumption. This method results in a lower bitrate after encoding and has only minimal adverse effects on the accuracy of machine vision tasks. In our experiments, we down-scale the luma channel of the input video using luma-scaling factors from 0.2 to 0.9 and evaluate coding results with optional back-scaling to the original range before machine vision tasks. Our results with the VVC Test Model (VTM) demonstrate that the proposed technique achieves coding gain of up to 37.9%and 46.1% for the same object detection and tracking accuracy, respectively.
Tero Partanen, Alban Marie, Alexandre Mercat, Jarno Vanne, Miska M. Hannuksela, Honglei Zhang 0001, Alireza Aminlou, Francesco Cricri
MMSP2
2024 Fast Machine Learning Aided Intra Mode Decision for Real-Time VVC Intra Coding
abstract
Reducing the huge computational complexity of intra mode decision is the key to real-time Video Coding (VVC). This paper proposes a fast intra mode decision scheme that takes advantage of lightweight machine learning (ML) models to classify intra modes into fifteen clusters. The cluster is further refined using one of the three proposed strategies to select the most optimal mode. Our experimental results with the fastest configuration of the practical uvg266 encoder show that the proposed methods yield a competitive rate-distortion-complexity trade-off over a conventional rough mode decision (RMD). To the best of our knowledge, this is the first work to successfully reduce the complexity of RMD in a practical VVC encoder with the use of ML techniques.
Joose Sainio, Baran Ataman, Alban Marie, Alexandre Mercat, Jarno Vanne
VCIP3
2023 Towards Machine Perception Aware Image Quality Assessment
abstract
Over the years, the objective of image and video compression has been to preserve perceived quality according to the Human Visual System (HVS) with minimal rate. Traditional encoders achieve this with the use of Rate-Distortion Optimization (RDO) techniques along with Image Quality Assessment (IQA) metrics that are correlated with human perception. Nowadays, a fast-growing number of applications fall within the realm of Video Coding for Machines (VCM), where the final recipient of compressed data is not a human but a machine performing a vision task. Recently, the lack of correlation between existing distortion measures and machine perception has been revealed, especially for RDO algorithms where distortion measures are computed on a local scale. In this paper, we propose a machine perception-aware metric designed to be incorporated into a standard-compliant Versatile Video Coding (VVC) encoder. Our proposed metric relies on a supervised training procedure as well as additional information available on the encoder side. In terms of correlation with machine perception, our metric significantly outperforms existing distortion measures in the literature.
Alban Marie, Karol Desnos, Jinjia Zhou, Luce Morin, Lu Zhang 0037
MMSP1
2023 Evaluation of Image Quality Assessment Metrics for Semantic Segmentation in a Machine-to-Machine Communication Scenario
abstract
Image and video compression aims at finding an optimal trade-off between rate and distortion. This is done through Rate-Distortion Optimization (RDO) in traditional en-coders with the use of Image Quality Assessment (IQA) metrics. While it is known that most IQA metrics are designed to be correlated with human perception, there is no evidence that this observation can be generalized in a Video Coding for Machines (VCM) context, where the receiver is not a human anymore but a machine. In this paper, we propose an evaluation protocol to measure the correlation level between conventional Full-Reference (FR) IQA metrics and machine perception through the semantic segmentation vision task. Experiments showed a relatively low correlation between them when measured on the block-level. This observation implies the need of RDO algorithms that are better suited for Machine-to-Machine (M2M) communications. In order to facilitate the emergence of IQA metrics that better reflect machine perception, the code and dataset used to perform this study is made freely available at https://github.com/albmarie/iqa_m2m_segmentation.
Alban Marie, Karol Desnos, Luce Morin, Lu Zhang 0037
QoMEX1
2022 Video Coding for Machines: Large-Scale Evaluation of Deep Neural Networks Robustness to Compression Artifacts for Semantic Segmentation
abstract
In the Video Coding for Machines (VCM) context where visual content is compressed before being transmitted to a vision task algorithm, appropriate trade-off between the compression level and the vision task performance must be chosen. In this paper, a Deep Neural Networks (DNN) based semantic segmentation algorithm robustness to compression artifacts is evaluated with a total of 1486 different coding configurations. Results indicate the importance of using an appropriate image resolution to overcome the block-partitioning limitations in existing compression algorithms, allowing 58.3%, 49.8%, 33.5% and 24.3% bitrate savings at equivalent prediction accuracy for JPEG, JM, x265 and VVenC, respectively. Surprisingly, JPEG can achieve 73.41% bitrate reduction with the inclusion of compressed images at training time over VVC Test Model (VTM) with a DNN trained on pristine data, which implies that DNN generalization ability must not be overlooked.
Alban Marie, Karol Desnos, Luce Morin, Lu Zhang 0037
MMSP1
2021 Rate-Distortion Optimized Motion Estimation for on-the-Sphere Compression of 360 Videos
abstract
On-the-sphere compression of omnidirectional videos is a very promising approach. First, it saves computational complexity as it avoids to project the sphere onto a 2D map, as classically done. Second, and more importantly, it allows to achieve a better rate-distortion tradeoff, since neither the visual data nor its domain of definition are distorted. In this paper, the on-the-sphere compression [1] for omnidirectional still images is extended to videos. We first propose a complete review of existing spherical motion models. Then we pro-pose a new one called tangent-linear+t. We finally propose a rate-distortion optimized algorithm to locally choose the best motion model for efficient motion estimation/compensation. For that purpose, we additionally propose a finer search pattern, called spherical-uniform, for the motion parameters, which leads to a more accurate block prediction. The novel algorithm leads to rate-distortion gains compared to methods based on a unique motion model.
Alban Marie, Navid Mahmoudian Bidgoli, Thomas Maugey, Aline Roumy
ICASSP1