Emmanuel Thomas

dblp:180/1704 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-8051-1091ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Depth map coding using depth range decomposition
abstract
Depth map sequences are commonly compressed using standard video codecs. However, the bitdepth of acquired depth data often exceeds the maximum bitdepth supported by deployed codecs, resulting in significant quantization loss. To overcome this, depth data are frequently mapped on multiple streams before encoding. In this paper, we propose a depth range decomposition of a high-bitdepth source depth map sequence into two lower-bitdepth sequences: a depth band mask sequence and a residual map sequence. The former indicates the depth band for each pixel, while the latter carries the residual within that band. The presented transformation is lossless and exploits the piecewise smoothness of depth data. Both sequences can be compressed independently with any standard video codec, allowing seamless integration into existing video processing pipelines.
Evangelos Alexiou, Emmanouil Potetsianakis, Emmanuel Thomas
VCIP3
2024 Management And Performance of Multiple Video Decoder Instances in Mobile Devices
abstract
Modern information exchange and telecommunication systems make immersive media increasingly more relevant. Immersive applications introduce challenges in ensuring optimal performance and scalability of media decoding, composition, and synchronization operations. Nowadays, a solution to accommodate media decoding in this context is to use multiple parallel video decoder instances instead of a single video decoder. This would require a robust and predictable decoder management system, coupled with dynamic buffer organization. In practice, however, this is often implemented without any special curation to manage the simultaneous video decoder instances, with resulting behaviour and implications not being fully understood yet. In this paper, we extend our previous work by utilizing the VidBench tool and investigate further the performance of modern chipsets integrated into Android-operated mobile devices in handling multiple video decoder instances running in parallel. Our results indicate the need for suitable synchronization solutions to counter uncertainty and variability across different video decoder instances and devices.
Evangelos Alexiou, Emmanouil Potetsianakis, Emmanuel Thomas
MMSys3
2024 Using Depth to Enhance Video-centric Applications
abstract
Acquiring depth data has become easily achievable with advancements in depth sensing and depth estimation technologies. As a result, obtaining a depth stream to describe the topology of a corresponding video stream has been considerably simplified. Presence of a depth stream offers numerous benefits, including the integration of advanced visual enhancements to the corresponding video stream in a flexible and efficient manner. This can enrich video-centric applications and facilitate their transition to Augmented Reality (AR) environments, where processing capabilities and battery power are limited. In this paper, we introduce VidDepth, an application developed for mobile devices to demonstrate examples of visual enhancements in video playback scenarios across both traditional and AR settings.
Emmanouil Potetsianakis, Evangelos Alexiou, Emmanuel Thomas, Emmanouil Xylakis
IMX3
2023 Video Decoding Performance and Requirements for XR Applications
abstract
Designing XR applications creates challenges regarding the performance and the scaling of media decoding operations, composition and synchronization of the various assets. Going beyond the single decoder paradigm of conventional video applications, XR applications tend to compose more and more visual streams such as 2D video assets but also textures and 2D/3D graphics encoded in video streams. All this demands a robust and predictable decoder management and a dynamic buffer organization. However, the behaviour of multiple decoder instances running in parallel is yet to be well understood on mobile platforms. To this end, we present in this paper VidBench - a parallel video decoding performance measurement tool for mobile Android devices. With VidBench, we quantify the challenges for applications using parallel video decoding pipelines with objective measurements and subjectively, we illustrate the current state of decoding multiple media streams and the possible visual artefacts resulting from unmanaged parallel video pipelines. Test results provide hints on the feasibility and the potential performance gain of using technologies like the MPEG-I Part 13 - Video Decoding Interface for immersive media (VDI) to alleviate those problems. We briefly present the main goals of VDI, standardised by the SC29 WG3 Moving Picture Experts Group (MPEG) Systems, which introduces functions and related constraints for optimizing such decoding instances as well as relevant video decoding APIs on which VDI is building upon such as the Khronos Vulkan Video extension.
Emmanouil Potetsianakis, Emmanuel Thomas
MMSys2
2023 MiroAR: Ubiquitous AR Teleconferencing Through The Mirror
abstract
Video call systems rely on being able to capture and transmit a self view, while at the same time rendering the view of the other party. Due to the lack of inwards facing cameras in XR devices (AR Glasses, HMD etc.) this is not a straightforward process. As a solution, recent XR teleconferencing platforms are trying to create a more "immersive" experience by replacing the self view with avatars, placing 3D models in space, creating shared spaces and other engaging features; approaches that are quite demanding and even then do not create a "traditional" teleconferencing experience. In this work, we are using an XR device (AR Glasses, or smartphone) to create a seamless and natural video calling experience. By using the AR Glasses to record an existing self-view from a reflective surface, like a mirror, the user is able to easily conduct a video call with a party, even if they are using a different setup. To demonstrate this concept, we present the MiroAR application. We conclude this paper by discussing the roadmap, shortcomings and possible extensions of our work.
Emmanouil Potetsianakis, Emmanuel Thomas
IMX2
2020 Fixed viewport applications for omnidirectional video content: combining traditional and 360 video for immersive experiences
abstract
With omnidirectional videos, the viewer is able to direct her Field-of-View (FoV) to any part of the scene while watching the content. This is achieved by rendering the 360 video content on the inside of a (conceptual) sphere in which the viewer is typically placed at the center. This is in contrast with traditional video that is rendered on a 2D plane and the viewer is watching always through a viewport directed by the content creator. These two approaches create a conflict between user experience and creativity, since omnidirectional video provides the user with viewing freedom, while traditional video allows for greater artistic expression by controlling the viewport. In order to combine these two approaches we propose an immersive setup in which the content changes between free-form viewing of omnidirectional (360 video mode) and directed viewing of traditional videos (director's mode). In this demo paper we present the benefits and reasoning behind this proposal and the means to implement it using the OMAF (MPEG-I - Part 2) standard.
Emmanouil Potetsianakis, Emmanuel Thomas, Karim El Assal, M. Oskar van Deventer
MMSys2
2020 VVC bitstream extraction and merging operations for multi-stream media applications
abstract
In traditional video decoding applications, the number of elementary streams that a hardware decoding platform of an end device can decode is determined at runtime by the. Upon request by the application, the decoding platform verifies whether a new decoding instance with an associated requirement in terms of data rate can fit under the current workload. Conversely, if a device can decode one 4K elementary stream in hardware, it may not be able to simultaneously decode four HD elementary streams that would each correspond to requirements in terms of data rate of 1/4 of the 4K elementary stream. Current video decoding platforms are thus designed with the assumption that each elementary stream requires the instantiation of a dedicated video decoder instance. At the same time, it has been increasingly common in new media applications such as immersive media applications to simultaneously consume several elementary streams in a synchronised fashion. The demo presents a new paradigm for media applications for which elementary streams may be consumed in such synchronised manner where the same decoder instance can be used. The demonstrator leverages on new features of the Versatile Video Coding (VVC) standard and interfaces being defined in the ongoing standardisation of MPEG-I part 13: Video Decoding Interface for Immersive Media. Stitching and cropping videos in the compressed domain can be achieved by an application via such defined interfaces. Without those interfaces, the same tasks are possible with the High Efficiency Video Coding (HEVC) standard to some extent but are tedious. In this demonstrator, we thus show how the new VVC codec can enable the decoupling of the number of elementary streams consumed by the application and the number of running video decoder instances. In addition, memory usage and CPU performance are also collected and compared with a tradition multiple decoding instance approach.
Emmanuel Thomas, Alexandre Gabriel, Karim El Assal
MMSys1
2019 Viewport-driven DASH media playback for interactive storytelling: a seamless non-linear storyline experience
abstract
Over the past few years, the HTTP Adaptive Streaming (HAS) technologies, e.g. the MPEG-DASH standard (DASH), became the predominant form of online video streaming. 360 Virtual Reality (VR) content have recently emerged on video streaming platform as a new immersive format but limits the interaction with the viewer to looking around via head rotation. From storytelling domain, interactive storytelling is a known technique to immerse the viewer into the story by allowing him to influence the way the story unfolds. In this demo, we combine a VR 360 video player and interactive storytelling delivered using DASH. By adding DASH-level signaling when branching occurs in the storyline and by updating the reference dash.js JavaScript player, it is possible to create an interactive 360 video experience that is seamless to the user. That is, each person experiences a different storyline based on its head movements whether it is by explicit cues, e.g. choosing a certain outcome to a scene, or implicitly by following narrative cues inserted by the content creator such as moving objects, character movements in the scene. Such seamless experience aims at maximizing the emotional response of the user on the content as intended by the content creator.
Karim El Assal, Emmanuel Thomas, Alexandre Gabriel, Sylvie Dijkstra-Soudarissanane
MMSys2
2018 Towards Low-Complexity Scalable Coding for Ultra-High Resolution Video And Beyond
abstract
The current state of the art of video coding lacks specific tools to efficiently deal with ultra-high resolutions such as 8K video both at a hardware and software level. At the same time, the increase in video resolutions combined with a larger variety in display resolutions results in a need for spatial scalability at low computational complexity. The proposed solution allows for legacy 4K encoders and decoders both software and in particular hardware to be still exploited with 8K content. The proposed format also offers spatial scalability for an average BD-rate increase of 16%. This scalability is achieved by a polyphase subsampling of the input sequence and by leveraging the existing temporal scalability to induce spatial scalability.
Emmanuel Thomas, Alexandre Gabriel, Omar Niamut, Sylvie Dijkstra-Soudarissanane
VCIP1
2016 Using MPEG DASH SRD for zoomable and navigable video
abstract
This paper presents a video streaming client implementation that makes use of the Spatial Relationship Description (SRD) feature of the MPEG-DASH standard, to provide a zoomable and navigable video to an end user. SRD allows a video streaming client to request spatial subparts of a particular video stream, which might be available in multiple resolutions.
Lucia D'Acunto, Jorrit van den Berg, Emmanuel Thomas, Omar Niamut
MMSys3
2016 MPEG DASH SRD: spatial relationship description
abstract
This paper presents the Spatial Representation Description (SRD) feature of the second amendment of MPEG DASH standard part 1, 23009-1:2014 [1]. SRD is an approach for streaming only spatial sub-parts of a video to display devices, in combination with the form of adaptive multi-rate streaming that is intrinsically supported by MPEG DASH. The SRD feature extends the Media Presentation Description (MPD) of MPEG DASH by describing spatial relationships between associated pieces of video content. This enables the DASH client to select and retrieve only those video streams at those resolutions that are relevant to the user experience. The paper describes the design principles behind SRD, the different possibilities it enables and examples of how SRD was used in different experiments on interactive streaming of ultra-high resolution video.
Omar Niamut, Emmanuel Thomas, Lucia D'Acunto, Cyril Concolato, Franck Denoual, Seong Yong Lim
MMSys2