Anthony Trioux

dblp:240/7361 · DBLP profile ↗
← Back
21ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0003-3457-3301ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 13 since 2021Computer networks · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Indoor Panoramic Depth Estimation via Dual-Projection Feature Fusion and Structural Refinement
Anthony Trioux, Yun Liang 0017
ISCAS3
2026 Avatar Standardization Efforts for Interoperable Metaverse Services: Toward a Seamless Avatar-as-a-Service Ecosystem
abstract
The metaverse has emerged as a transformative platform that blends physical and digital realities, supporting immersive communication, interaction, and service delivery. At the core of these experiences are avatars, digital embodiments of users, that serve as interactive interfaces across a wide range of metaverse services, including virtual collaboration, education, healthcare, entertainment, e-commerce,etc.However, the fragmented ecosystem of tools, platforms, and proprietary formats has led to significant challenges in avatar interoperability, preventing seamless user identity and experience across applications. In response, numerous international standardization bodies have launched initiatives to define interoperable and modular avatar frameworks. This survey represents the first comprehensive review of these initiatives with a focus on service-oriented requirements, revealing the concept of Avatar-as-a-Service (AaaS). It examines both historical foundations and recent advances, with a particular attention given to the emerging ISO/IEC 23090-39 standard (MPEG-I Part 39 Avatar Representation Format, ARF). By synthesizing technical approaches, standardization efforts, and open challenges, this work presents a roadmap for the evolution of AaaS and highlights emerging opportunities for delivering personalized, high-fidelity, and cross-compatible avatar experiences to support the next generation of metaverse services.
Anthony Trioux, Wei Zhang 0072, Giuseppe Valenzise, Fuzheng Yang 0001
IEEE Trans. Serv. Comput.1
2025 Blendshape Compression Techniques and Their Impact on Reconstructed Avatar Face Animation: A Subjective Study
abstract
Blendshapes have been widely adopted as a key method for generating facial animation on avatars due to their ease of manipulation, flexibility in capturing diverse facial expressions, and compatibility with real-time rendering. However, current frameworks lack efficient methods for compressing blendshape (BS) animation parameters, which are critical for optimizing data transmission. This study introduces a pioneer compression scheme leveraging the amount of BS to be transmitted, their quantization as well as their transmission frequency. Subjective evaluation using the ITU-R BT.500-15 recommendation demonstrates that the proposed method significantly reduces the amount of data to transmit while preserving acceptable visual quality. This approach extends prior findings on reduced BS sets[kang2023effects] and addresses a significant gap in avatar media coding[avril2023morgan]. This work establishes a baseline for efficient facial animation and, serves as a foundation for further exploration on adaptive rate-allocation strategies and advanced compression strategies tailored for diverse avatar animation scenarios.
Anthony Trioux, Wei Zhang 0072, Yusong Gao, Giuseppe Valenzise, Fuzheng Yang 0001
DCC1
2025 Subjective Fidelity Assessment of Audio- and Video-Driven Talking Head Generation Methods
abstract
Audio- and Video-Driven Talking Head Generation methods have attracted considerable research interest due to recent advances in Artificial Intelligence Generated Content (AIGC) technologies. In such approaches, a single image is artificially animated by leveraging audio and/or motion features extracted from video sources. Despite notable progress, current performance assessments rely primarily on traditional objective metrics, often neglecting subjective evaluation aspects. To address this issue, we propose in this paper a subjective fidelity assessment of recent Audio- and/or Video-Driven Talking Head Generation methods. This study aims to assess how accurately and convincingly the generated video reproduces the visual and behavioral characteristics of a real human face, as well as how closely the video aligns with expected natural human expressions, movements, and/or audio synchronization. In order to provide a detailed assessment of the fidelity in the context of talking heads, our study focuses on six key criteria: Overall Fidelity, Gaze Fidelity, Audio-Video Sync Fidelity, Head Pose Fidelity, Expression Fidelity, and Overall Visual Quality. Experiments results reveal a nuanced picture of the fidelity in this context, where the performance varies significantly depending on the video content itself as well as how the animation is generated, highlighting the needs for further research. This research represents an initial step towards the evaluation of Audio- and Video-Driven generative image animation methods for Talking heads while offering insights for improving the accuracy and realism of those techniques. The dataset and corresponding results are available at https://github.com/a-trioux/Subjective-Fidelity-Assessment-Talking-Head.
Anthony Trioux, Yusong Gao, Jiarun Song, Faming Ma, Fuzheng Yang 0001
ICASSP1
2025 Exploring Compression Strategies for Blendshape-Based Avatar Facial Animation: Subjective and Objective Analysis
abstract
Blendshapes (BS) have been widely adopted as a key method for generating facial animations on avatars due to their ease of manipulation, flexibility in capturing diverse expressions, and compatibility with real-time rendering. Despite their importance, current frameworks, including the recent MPEG Avatar Representation Format one, lack efficient methods for blendshape compression. This paper addresses this gap by proposing and analyzing a pioneering compression scheme for blendshape animation parameters. The method explores three key dimensions of compression: the number of transmitted BS, their quantization, as well as their frequency of transmission. The use of a linear interpolation for reduced blendshape transmission rates allow to mitigate flickering introduced by strong quantization, enhancing the viewing experience. Subjective evaluations demonstrate that the proposed approach achieves substantial data transmission savings while maintaining acceptable visual quality. Furthermore, an in-depth comparison of classical objective metrics against Mean Opinion Scores (MOS) reveals their limitations in accurately capturing perceived quality after blendshape compression, with Pearson and Spearman correlation scores reaching at most ~ 0.6. Among the evaluated metrics, Detail Loss Metric (DLM), Video Multimethod Assessment Fusion (VMAF), and Mean Peak Signal-to-Noise Ratio in the BS domain (PSNR-BS) exhibit the highest correlation with MOS. This study provides a comprehensive benchmark for blendshape compression and has driven the creation of an Exploratory Experiment (EE) within the ongoing MPEG avatar-related efforts, highlighting the study’s relevance to standardization activities.
Anthony Trioux, Wei Zhang 0072, Giuseppe Valenzise, Fuzheng Yang 0001
ICME1
2025 Subjective Visual Quality Assessment of Compressed Light Field Images: Learning-based vs. Conventional Methods
abstract
Light fields (LF) technology enables the capture and reproduction of a 3D scene accurately, which enhances visual experience in various applications. The sheer volume of multiplicity of the captured views creates logistical problems in both storage and data transmission, which makes LF compression crucial. Even though many LF compression techniques have been proposed and evaluated in recent years, the leading-edge learning-based approaches have not been subject to the same level of scrutiny. This paper presents a subjective quality assessment study on four different LF compression methods, including two learning-based LF compression methods, which have not been studied before from a subjective quality point of view. For this purpose, subjective opinion scores were collected from viewers in two different universities for a cross-lab study. The results indicate that the learning-based compression methods have different behavior in their rate-distortion curves, and that there is room for improvement for learning-based methods. A qualitative analysis also shows that their artifact structures are different from conventional ones. The results highlight the need for a perceptual objective quality metric that takes different types of artifacts into account. The obtained subjective quality database (MiX-LFQDB) is made public to support further research in this area: https://doi.org/10.5281/zenodo.16778670
Emin Zerman, Soheib Takhtardeshir, Anthony Trioux, Jianlong Qin, Roger Olsson, Mårten Sjöström
MMSP3
2025 CV-Cast: Computer Vision-Oriented Linear Coding and Transmission
abstract
Remote inference allows lightweight edge devices, such as autonomous drones, to perform vision tasks exceeding their computational, energy, or processing delay budget. In such applications, reliable transmission of information is challenging due to high variations of channel quality. Traditional approaches involving spatio-temporal transforms, quantization, and entropy coding followed by digital transmission may be affected by a sudden decrease in quality (thedigital cliff) when the channel quality is less than expected during design. This problem can be addressed by using Linear Coding and Transmission (LCT), a joint source and channel coding scheme relying on linear operators only, allowing to achieve reconstructed per-pixel error commensurate with the wireless channel quality. In this paper, we propose CV-Cast: the first LCT scheme optimized for computer vision task accuracy instead of per-pixel distortion. Using this approach, for instance at 10 dB channel signal-to-noise ratio, CV-Cast requires transmitting 28% less symbols than a baseline LCT scheme in semantic segmentation and 15% in object detection tasks. Simulations involving a realistic 5G channel model confirm the smooth decrease in accuracy achieved with CV-Cast, while images encoded by JPEG or learned image coding (LIC) and transmitted using classical schemes at low Eb/N0 are subject to digital cliff.
Jakub Zádník, Michel Kieffer, Anthony Trioux, Markku Mäkitalo, Pekka Jääskeläinen
IEEE Trans. Mob. Comput.3
2025 Correction to "CV-Cast: Computer Vision-Oriented Linear Coding and Transmission"
abstract
In the above article [1], on page 1151, eq. (6), there is an error in the equation. The correct equation is: \begin{equation*} \min.\,\,D,\,\,\text{s.t.} \sum\limits_{k = 1}^K {{{\lambda }_k}\beta _k^2 \leqslant P.} \tag{6} \end{equation*} min.D,s.t.∑k=1Kλkβk2⩽P.(6)
Jakub Zádník, Michel Kieffer, Anthony Trioux, Markku Mäkitalo, Pekka Jääskeläinen
IEEE Trans. Mob. Comput.3
2024 Improving Reconstruction Fidelity in Generative Face Video Coding using High-Frequency Shuttling
abstract
Generative face video coding (GFVC) schemes applied to talking head videos have recently demonstrated significant coding gains compared to traditional coding frameworks, particularly at ultra-low bitrates. Despite advancements in the field, these methods still face challenges in handling large pose and facial expression changes, as well as (dis-)occlusions. Recently, a hybrid approach (HDAC+) that combines a low-quality video coded with a conventional codec and animation-based coding has been proposed for standardization and shown to partially address these issues. Although HDAC+ shows promising results, it still struggles with generating accurate images. In this paper, we propose HDAC-HF, an improvement to the reconstruction process in HDAC+. Based on empirical observations that some pose and expression details are lost during animation, we introduce a high-frequency (HF) shuttling mechanism to enhance reconstruction fidelity at the decoder side, inspired by recent advancements in video super-resolution. By enhancing the flow of high-frequency details in the feature domain, we improve the reconstruction of facial expressions and poses. Qualitative and quantitative experiments confirm that the proposed method improves reconstruction without any additional bitstream or signaling cost compared to the baseline HDAC+ codec.
Goluck Konuko, Giuseppe Valenzise, Anthony Trioux
VCIP3
2024 Performance of Linear Coding and Transmission in Low-Latency Computer Vision Offloading
abstract
Image communication increasingly involves machine-to-machine delivery. For example, images acquired by an autonomous drone can be compressed and sent to an edge server over a wireless network for resource-intensive processing. Traditional compression techniques involving transform, quantization, and entropy coding reach high compression efficiency, but channel conditions worse than expected may lead to a sharp decrease in the decoded image quality. As an alternative, Linear Coding and Transmission (LCT) systems have been proposed to avoid this digital cliff problem: The reconstructed image quality decreases gradually as channel conditions degrade. This paper presents a comprehensive evaluation of computer vision tasks with input images processed and transmitted using LCT. It also analyses the benefits of network retraining, accounting for impairments due to LCT and noisy channel. Considering object detection and semantic segmentation over images transmitted and received by LCT systems, we show that the task accuracy degrades smoothly when the channel quality decreases, avoiding the cliff effect. Retraining with noisy images processed by LCT restores detection mAP degradation from 23.8% to 4.4% and segmentation mIoU degradation from 43.2% to 8.1 % when the channel signal-to-noise ratio is 10 dB.
Jakub Zádník, Anthony Trioux, Michel Kieffer, Markku Mäkitalo, François-Xavier Coudoux, Patrick Corlay, Pekka Jääskeläinen
WCNC2
2023 Glass-to-Glass Delay Reduction: Encoding Rate Reduction vs. Video Frame Extrapolation
abstract
Applications such as teleoperated driving, remote robot control, and telepresence rely on video services to ensure real-time interaction with a satisfying quality of experience. Reducing the Glass-to-Glass (G2G) delay, i.e., the time delay between the acquisition of a video frame and its display on a remote terminal is critical for these applications. Deep learning-based video frame extrapolation before video encoding has been recently considered as an interesting solution to reduce G2G delay, however, the latency introduced by extrapolation has not been taken into account. In this paper, considering the main sources of latency, including extrapolation delay, we examine the benefits and limitations of frame extrapolation at encoder in reducing the G2G delay in a point-to-point video transmission system. To this end, we compare the latency-quality trade-off for two latency compensation methods: encoding rate reduction and video frame extrapolation. Our aim is to determine the G2G delay reduction that may be achieved at the price of a given quality reduction. Our experiments show that extrapolation methods can provide a null perceived G2G delay with an acceptable loss in quality, particularly for applications with video contents with limited temporal information. Such delay reduction is unreachable via encoding rate reduction.
Hind Kanj, Anthony Trioux, Marco Cagnazzo, François-Xavier Coudoux, Patrick Corlay, Michel Kieffer
MMSP2
2023 Effect of latency on social presence in traditional video conference and VR conference: a comparative study
abstract
Virtual reality (VR) conference, as a typical social VR application, has gained popularity in recent years. It offers users located at different locations a fully immersive experience and a sense of togetherness. However, the remote communication also introduces inevitable latencies, which may adversely affect the so-called social presence. There is still a lack of research on the effect of latency on social presence. To fill the gap, this paper aims to examine the impact of latency on social presence of VR conference and contrast it with that of traditional video conference. Here, the social presence is measured using the Networked Minds Social Presence Inventory (NMSPI). We design and conduct two conversation-based subjective tests for both types of conference and compare the impact of the latency based on the test results. The conclusions of these studies can be used as guidelines for VR service providers to optimize their conference systems.
Jiarun Song, Anthony Trioux, Yusong Gao, Fuzheng Yang 0001
VCIP3
2023 An efficient video-based geometry compression system for 3D meshes
abstract
The ISO/IEC JTC1 SC29 subtechnical committee including the Moving Picture Experts Group (MPEG) is currently working on a Video-based Dynamic Mesh Coding (V-DMC) standard to enable next-generation video applications. Relying on conventional 2D video-based solutions for such content allows to take advantage of the maturity of current video codecs. However, directly re-use such codecs is inefficient as images projected from sparse 3D meshes are different from classical video content. In this paper, we present a video-based system to compress the geometry information of sparse 3D meshes in a highly efficient way. Specifically, inspired by the Visual Volumetric Video-based Coding standard (V3C) standard, we propose a method to encode meshes by using orthogonal projections, followed by atlas packing and 2D video coding. The proposed method includes a quantization step that allows to improve the video coding efficiency while being compatible with lossless and lossy compression. This ability is highly encouraged by the V-DMC standard. Experimental results on the V-DMC test sequences demonstrate that the proposed method, leveraging well-known conventional 2D video codec, obtains a high coding efficiency for compressing the geometry information of 3D meshes. Specifically, an average bitrate saving of 64.58% for the lossless case is observed, and an average bitrate saving of 58.82% and BD-rate gain of about -58% for the lossy case are obtained, respectively.
Wenjie Zou, Haidi Huang, Anthony Trioux, Fuzheng Yang 0001
VCIP3
2023 Mmetric++: An improved quality metric to judge the lossless compression of 3D meshes
abstract
During the past decade, 3D meshes have been widely used to represent immersive content The Moving Picture Experts Group (MPEG) is working on a new standard called Video-based Dynamic Mesh Coding (V-DMC) for compressing dynamic meshes. Due to market demand, lossless compression represents an important test condition for this standard. The quality metric denoted mmetric currently used by the standard to verify the lossless character of a proposed codec evaluates the geometry and attributes coordinates differences. However, 3D meshes are also characterized by their connectivity, which is a parameter completely overlooked by the current metric, leading to wrong decisions when evaluating the lossless character of a 3D mesh codec. In this paper, we propose an improved quality metric to correctly judge whether a mesh has been losslessly encoded or not. The proposed method employs the configurations of the triangle fan (TF) defined in Triangle Fan-based compression (TFAN) to traverse the original mesh and the decoded one in order to compare their connectivity. Experimental results on the V-DMC test sequences demonstrate that the proposed method can effectively determine the differences existing in the connectivity. The proposed metric has been recently adopted and will be integrated soon into the new version of the mmetric tool set.
Wenjie Zou, Xinhang He, Anthony Trioux, Fuzheng Yang 0001
VCIP3
2023 Deep learning assisted quality ranking for list decoding of videos subject to transmission errors
abstract
In this paper, we propose a new deep learning-based quality ranking framework to assist video list decoding methods in the context of unreliable video transmissions. The objective is to identify an intact image (corrected video frame) among a list of candidate images generated by a list decoding method, where all candidates, except for the intact image are corrupted. The framework comprises a deep learning-based no-reference image quality assessment (NR-IQA) for non-uniform video distortions (NUD) system to rank the candidate images according to their quality, which allows identifying the best one. To show the validity of our proposed framework, we develop an NR-IQA system relying on a proven patch-based convolutional neural network (CNN) architecture, which we adapt to better account for the non-uniform distortions observed in the candidate images, e.g., H.265 transmission errors during wireless communications. Specifically, we modify the patch size on which our CNN for non-uniform distortions (CNN-NUD) operates to capture a larger and more meaningful spatial context. Moreover, we develop a new training database using images resulting from various bit modifications in the received video packets, to simulate the list decoding process, and train the system using a full reference IQA (FR-IQA) method. Experiments on intra frames of videos encoded using H.265 show the ability of this system to identify an intact image among a set of five candidate images with an average accuracy of 96.6%, whereas traditional NR-IQA metrics or the initially trained CNN system offer poor accuracy ranging between 15.7% and 33.6%, respectively.
Alexis Guichemerre, Stéphane Coulombe, Anthony Trioux, François-Xavier Coudoux, Patrick Corlay
WiMob3
2022 Towards Zero-Latency Video Transmission Through Frame Extrapolation
abstract
In the past few years, several efforts have been devoted to reduce individual sources of latency in video delivery, including acquisition, coding and network transmission. The goal is to improve the quality of experience in applications requiring real-time interaction. Nevertheless, these efforts are fundamentally constrained by technological and physical limits. In this paper, we investigate a radically different approach that can arbitrarily reduce the overall latency by means of video extrapolation. We propose two latency compensation schemes where video extrapolation is performed either at the encoder or at the decoder side. Since a loss of fidelity is the price to pay for compensating latency arbitrarily, we study the latency-fidelity compromise using three recent video prediction schemes. Our preliminary results show that by accepting a quality loss, we can compensate a typical latency of 100 ms with a loss of 8 dB in PSNR with the best extrapolator. This approach is promising but also suggests that further work should be done in video prediction to pursue zero-latency video transmission.
Melan Vijayaratnam, Marco Cagnazzo, Giuseppe Valenzise, Anthony Trioux, Michel Kieffer
ICIP4
2021 A Perceptual Study of the Decoding Process of the SoftCast Wireless Video Broadcast Scheme
abstract
The SoftCast scheme has been proposed as a promising alternative to traditional video broadcasting systems in wireless environments. In its current form, SoftCast performs image decoding at the receiver side by using a Linear Least Square Error (LLSE) estimator. Such approach maximizes the reconstructed quality in terms of Peak Signal-to-Noise Ratio (PSNR). However, we show that the LLSE induces an annoying blur effect at low Channel Signal-to-Noise Ratio (CSNR) quality. To cancel this artifact, we propose to replace the LLSE estimator by the Zero-Forcing (ZF) one. In order to better understand the perceived quality offered by these two estimators, a mathematical characterization as well as an objective and subjective studies are performed. Results show that the gains brought by the LLSE estimator, in terms of PSNR and Structural SIMiliraty (SSIM), are limited and quickly tend to null value as the CSNR increases. However, higher gains are obtained by the ZF estimator when considering the recent Video Multi-method Assessment Fusion (VMAF) metric proposed by Netflix, which evaluates the perceptual video quality. This result is confirmed by the subjective assessment.
Anthony Trioux, Giuseppe Valenzise, Marco Cagnazzo, Michel Kieffer, François-Xavier Coudoux, Patrick Corlay, Mohamed Gharbi
MMSP1
2021 CRC-Based Multi-Error Correction of H.265 Encoded Videos in Wireless Communications
abstract
This paper analyzes the benefits of extending CRC-based error correction (CRC-EC) to handle more errors in the context of error-prone wireless networks. In the literature, CRC-EC has been used to correct up to 3 binary errors per packet. We first present a theoretical analysis of the CRC-EC candidate list while increasing the number of errors considered. We then analyze the candidate list reduction resulting from subsequent checksum validation and video decoding steps. Simulations conducted on two wireless networks show that the network considered has a huge impact on CRC-EC performance. Over a Bluetooth low energy (BLE) channel with Eb/No=8 dB, an average PSNR improvement of 4.4 dB on videos is achieved when CRC-EC corrects up to 5, rather than 3 errors per packet.
Vivien Boussard, Stéphane Coulombe, François-Xavier Coudoux, Patrick Corlay, Anthony Trioux
VCIP5
2021 A comprehensive theoretical evaluation of the end-to-end performance of SoftCast-based linear video delivery schemes
Anthony Trioux, Mohamed Gharbi, François-Xavier Coudoux, Patrick Corlay
Signal Process. Image Commun.1
2020 Subjective and Objective Quality Assessment of the SoftCast Video Transmission Scheme
abstract
SoftCast-based linear video coding and transmission (LVCT) schemes have been proposed as a promising alternative to traditional video coding and transmission schemes in wireless environments. Currently, the performance of LVCT schemes is evaluated by means of traditional objective scores such as PSNR or SSIM. Nevertheless, since the compression is performed in a very different way from traditional coding schemes such as HEVC, visual artifacts are also quite different and deserve to be subjectively assessed. In this paper, we propose a subjective quality assessment of SoftCast, pioneer and standard of the LVCT schemes. This study aims to better understand the trade-offs between the LVCT parameters that can be tuned to improve the quality. These parameters, including different GoP-sizes, Compression Ratios (CR) and Channel Signal-to-Noise Ratio (CSNR), are used to generate a dataset of 85 videos. A Double Stimulus Impairment Scale (DSIS) test is performed on the received videos to assess the perceived quality. Results show that the key characteristic of SoftCast, the linear relation between CSNR and PSNR, is also observed with the Mean-Opinion Scores (MOS), except at high CSNR where the quality saturates. In addition, Bjøntegaard model is used to quantify the trade-offs between CR, GoP-size and CSNR, depending on the intended application. Finally, the performance of objective metrics compared to the obtained MOS is evaluated. Results show that Multi-Scale SSIM (MS-SSIM), SSIM and Video Multimethod Assessment Fusion (VMAF) metrics offer the best correlation with the MOS values.
Anthony Trioux, Giuseppe Valenzise, Marco Cagnazzo, Michel Kieffer, François-Xavier Coudoux, Patrick Corlay, Mohamed Gharbi
VCIP1
2020 Temporal information based GoP adaptation for linear video delivery schemes
Anthony Trioux, François-Xavier Coudoux, Patrick Corlay, Mohamed Gharbi
Signal Process. Image Commun.1