Omar Niamut

dblp:122/0974 · also Omar Aziz Niamut · DBLP profile ↗
← Back
25ranked-venue papers
6as first author
9since 2021 · last 2024
0000-0002-2398-0465ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 6 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 since 2021
YearPublicationVenuePosition
2024 VP9 bitstream-based Tiled Multipoint Control Unit: Scaling simultaneous RGBD user streams in an immersive 3D communication system
abstract
Video conference applications that allow group communication through video and audio over distance have become commonplace and mainstream. To scale the number of participants in video conferencing systems, the usual practice is to deploy centralized streaming components such as Multipoint Control Units (MCU) or Selective Forwarding Units (SFU). Using these components can become problematic due to significant resource overhead in the server or client. In this paper, we propose a Tile-aware Multipoint Control Unit (T-MCU) that is capable of combining and forwarding streams on a bitstream level. In our solution, multiple VP9 user video streams are combined into a single bitstream, requiring only a single video transmission (single receiving socket, single video decoder, and single rendering texture). Multiple changes (relating to dynamic header and motion vectors) had to be made to the reference VP9 encoder to allow for combining bitstream tiles. These changes are available as open source and are described in this paper, including their impact on encoding performance. Furthermore, all changes are fully compatible with existing VP9 decoders. The paper also presents an experimental design study to evaluate our new T-MCU under different streaming conditions (including the impact on the receiving client) within an immersive 3D communication system utilizing RGBD video data. Our approach significantly reduces the performance requirements of the client (10-20%) and server (~80%) at the cost of increased bandwidth (~16%). Ultimately, the T-MCU allows immersive 3D communication applications to support at least 16+ simultaneous users.
Simon Gunkel, Rick Hindriks, Yonatan Shiferaw, Sylvie Dijkstra-Soudarissanane, Omar Niamut
MMSys5
2023 "You AR' right in front of me": RGBD-based capture and rendering for remote training
abstract
Immersive technologies such as virtual reality have enabled novel forms of education and training, where students can learn new skills in simulated environments. But some specialized training procedures, e.g. ESA-certified soldering, still involve real-world physical processes with physical lab equipment. Such training sessions require students to travel to teaching labs and may interrupt everyday commitments for a longer period of time. There is a desire to make such training procedures more accessible remotely while keeping any student-to-teacher interaction natural, personal, and engaging. This paper presents a prototype for a remote teaching use case by rendering 3D photorealistic representations into the Augmented Reality (AR) glasses of a student. The teacher is captured with a modular RGBD capture application integrated into a web-based immersive communication platform. The integration offers multiple real-time capture calibration and rendering configurations. Our modular platform allows for an easy evaluation of different technical constraints as well as easy testing of the use case itself. Such evaluation may include a direct comparison of different 3D point-cloud and mesh rendering techniques. Additionally, the overall system allows immersive interaction between the student and the teacher, including augmented text messages for non-intrusive notifications. Our platform offers an ideal testbed for both technical and user-centered immersive communication studies.
Simon Gunkel, Sylvie Dijkstra-Soudarissanane, Omar Niamut
MMSys3
2023 Remote Expert Assistance System for Mixed-HMD Clients over 5G Infrastructure
abstract
When operating under adverse conditions or at distant locations, it is not always feasible to obtain the assistance of an expert on site. In such cases, remote expert assistance may provide a solution. Current remote assistance systems employ immersive data visualization or video-based communication. By introducing XR-based multi-user communication system to remote expert assistance solutions, we expect to improve the effectiveness of both the operator and the supporting experts. Our multi-user XR collaboration demo integrates XR and cloud/network technologies. It enables three users to collaborate in a way that makes them feel that they are solving a challenging task together.
Frank Bart ter Haar, Sylvie Dijkstra-Soudarissanane, Piotr Zuraniewski, Rick Hindriks, Karim El Assal, Simon Gunkel, Galit Rahim, Omar Niamut
MMSys8
2022 Deep Learning Augmented Realistic Avatars for Social VR Human Representation
abstract
Virtual reality (VR) has created a new and rich medium for people to meet each other digitally. In VR, people can choose from a broad range of representations. In several cases, it is important to provide users with avatars that are a lifelike representation of themselves, to increase the user experience and effectiveness of communication. In this work, we propose a pipeline for generating a realistic and expressive avatar from a single reference image. The pipeline consists of a blendshape-based avatar combined with two deep learning improvements. The first improvement module runs offline and improves the texture map of the base avatar. The second module runs inference in real-time at the rendering stage and performs a style transfer to the avatar’s eyes. The deep learning modules effectively improve the visual representation of the avatar and show how AI techniques can be integrated with traditional animation methods to generate realistic human avatars for social VR.
Matthijs van der Boon, Leonor Fermoselle, Frank Bart ter Haar, Sylvie Dijkstra-Soudarissanane, Omar Niamut
IMX5
2021 Dynamic Edge Offloading for Real-time Video Processing Pipelines
abstract
In this demo, we show a system where real-time processing modules for immersive conferencing streams can be dynamically relocated between user device and edge computing nodes, with minimal visual impact on the resulting stream.
Jan Willem Kleinrouweler, Toni Dimitrovski, Sjors Braam, Rick Hindriks, Hans van den Berg, Lucia D'Acunto, Omar Niamut
MMSP7
2021 XR Carousel: A Visualization Tool For Volumetric Video
abstract
Recent years have seen a new uptake in immersive media and eXtended Reality (XR). And due to a global pandemic, computer-mediated communication over video conferencing tools became a new normal of everyday remote collaboration and virtual meetings. Social XR leverages XR technologies for remote communication and collaboration. But in order for XR to facilitate a high level of (social) presence and thus high-quality mediated social contact between users, we need high-quality 3D representation of users. One approach to providing detailed 3D user representations as new immersive media is to use point clouds or meshes, but these representation formats come with complexity on compression bitrate and processing time. In the example of virtual meetings, compression has to fulfill stringent requirements such as low latency and high quality. As the compression techniques for 3D immersive media steadily advance, it is important to be able to easily compare different compression techniques on their technical and visual merits in an easy way. The proposed demonstrator in this paper is a visualization tool that helps assessing the visual quality of a 3D representation employing various coding schemes. The complete end-to-end rendering/encoding chain can be easily assessed, allowing for subjective testing by showing the differences between the selected encoding parameters. The tool presented in this demo paper offers an improved and easy visual process for the comparison of encoders of immersive media.
Sylvie Dijkstra-Soudarissanane, Simon Gunkel, Alexandre Gabriel, Leonor Fermoselle, Frank Bart ter Haar, Omar Niamut
MMSys6
2021 VRComm: an end-to-end web system for real-time photorealistic social VR communication
abstract
Tools and platforms that enable remote communication and collaboration provide a strong contribution to societal challenges. Virtual meetings and conferencing, in particular, can help to reduce commutes and lower our ecological footprint, and can alleviate physical distancing measures in case of global pandemics. In this paper, we outline how to bridge the gap between common video conferencing systems and emerging social VR platforms to allow immersive communication in Virtual Reality (VR). We present a novel VR communication framework that enables remote communication in virtual environments with real-time photorealistic user representation based on colour-and-depth (RGBD) cameras and web browser clients, deployed on common off-the-shelf hardware devices. The paper's main contribution is threefold: (a) a new VR communication framework, (b) a novel approach for real-time depth data transmitting as a 2D grayscale for 3D user representation, including a central MCU-based approach for this new format and (c) a technical evaluation of the system with respect to processing delay, CPU and GPU usage.
Simon Gunkel, Rick Hindriks, Karim El Assal, Hans Stokking, Sylvie Dijkstra-Soudarissanane, Frank Bart ter Haar, Omar Niamut
MMSys7
2021 Towards XR Communication for Visiting Elderly at Nursing Homes
abstract
Due to the current pandemic, the elderly in care homes are greatly affected by the lack of contact with their families, resulting in various mental conditions (e.g., depression, feelings of loneliness) and deterioration of mental health for dementia patients. In response, residents and family members increasingly resorted to mediated communication to maintain social contact. To facilitate high-quality mediated social contact between residents in nursing homes and remote family members, we developed an Augmented Reality (AR)-based communication tool. The proposed demonstrator improved this situation by providing a working communication tool that enables the elderly to feel being together with their family by means of AR techniques. A complete end-to-end-chain architecture is defined, where the aspects of capture, transmission, and rendering are thoroughly investigated to fit the purpose of the use case. Based on an extensive user study comprising user experience (UX) and quality of service (QoS) measurements, each module is presented with the improvements made and the resulting higher quality AR communication platform.
Sylvie Dijkstra-Soudarissanane, Tessa Klunder, Aschwin Brandt, Omar Niamut
IMX4
2021 Augmented Reality-Based Remote Family Visits in Nursing Homes
abstract
During the COVID-19 pandemic, many nursing homes had to restrict visitations. This had a major negative impact on the wellbeing of residents and their family members. In response, residents and family members increasingly resorted to mediated communication to maintain social contact. To facilitate high-quality mediated social contact between residents in nursing homes and remote family members, we developed an augmented reality (AR)-based communication tool. In this study, we compared the user experience (UX) of AR-communication with that of video calling, for 10 pairs of residents and family members. We measured enjoyment, spatial presence and social presence, attitudes, behavior and conversation duration. In the AR-communication condition, residents perceived a 3D projection of their remote family member onto a chair placed in front of them. In the video calling condition, the family member was shown using 2D video. In both conditions, the family member perceived the resident in the video calling mode on a 2D screen. While residents reported no differences in their UX between both conditions, family members reported higher spatial presence for the AR-communication condition compared to video-calling. Conversation durations were significantly longer during AR-communication than during video calling. We tentatively suggest that there may be (unconscious) differences in UX during AR-based communication compared to video calling.
Alexander Toet, Hans Stokking, Tessa Klunder, Zeph M. C. van Berlo, Bram Smeets, Omar Niamut
IMX6
2020 Let's Get in Touch! Adding Haptics to Social VR
abstract
Social VR shall allow natural communication between users with high social presence, as if users are in the same room. One way to increase social presence is to add haptic interaction to allow, for example, users to give each other a ”high-five” or to pass documents among them. In this paper, we present our web-based VR communication framework with an added haptic component to simulate touch. The goal of this framework is to enhance the VR communication experience and the social cues exchange between users in VR. We describe our method for rendering haptic feedback within the web-based framework and evaluate the perceived quality of our system with a user survey (with 119 participants). Our proof-of-concept system was rated positively, with the haptic component offering an enhanced quality of the VR experience for 78% of the participants.
Leonor Fermoselle, Simon Gunkel, Frank Bart ter Haar, Sylvie Dijkstra-Soudarissanane, Alexander Toet, Omar Niamut, Nanda van der Stap
IMX6
2019 Multi-sensor capture and network processing for virtual reality conferencing
abstract
Recent developments in key technologies like 5G, Augmented and Virtual Reality (AR/VR) and Tactile Internet result in new possibilities for communication. Particularly, these key digital technologies can enable remote communication and collaboration in remote experiences. In this demo, we work towards 6-degrees of freedom (DoF) photo-realistic shared experiences by introducing a multi-view multi-sensor capture end-to-end system. Our system acts as a baseline end-to-end system for capture, transmission and rendering of volumetric video of user representations. To handle multi-view video processing in a scalable way, we introduce a Multi-point Control Unit (MCU) to shift processing from end devices into the cloud. MCUs are commonly used to bridge videoconferencing connections, and we design and deploy a VR-ready MCU to reduce both upload bandwidth and end-device processing requirements. In our demo, we focus on a remote meeting use case where multiple people can sit around a table to communicate in a shared VR environment.
Sylvie Dijkstra-Soudarissanane, Karim El Assal, Simon Gunkel, Frank Bart ter Haar, Rick Hindriks, Jan Willem Kleinrouweler, Omar Niamut
MMSys7
2019 360-Degree Photo-realistic VR Conferencing
abstract
VR experiences are becoming more social, but many social VR systems represent users as artificial avatars. For use cases such as VR conferencing, photo-realistic representations may be preferred. In this paper, we present ongoing research into social VR experiences with photo-realistic representations of participants and present a web-based social VR framework that extends current video conferencing capabilities with new VR functionalities. We explain the underlying design concepts of our framework and discuss user studies to evaluate the framework in three different scenarios. We show that people are able to use VR communication in real meeting situations and outline our future research to better understand the actual benefits and limitations of our approach, to fully understand the technological gaps that need to be bridged and to better understand the user experience.
Simon Gunkel, Marleen D. W. Dohmen, Hans Stokking, Omar Niamut
VR4
2018 Virtual reality conferencing: multi-user immersive VR experiences on the web
abstract
Virtual Reality (VR) and 360-degree video are set to become part of the future social environment, enriching and enhancing the way we share experiences and collaborate remotely. While Social VR applications are getting more momentum, most services regarding Social VR focus on animated avatars. In this demo, we present our efforts towards Social VR services based on photo-realistic video recordings. In this demo paper, we focus on two parts, the communication between multiple people (max 3) and the integration of new media formats to represent users as 3D point clouds. We enhance a green screen (chroma key) like cut-out of the person with depth data, allowing point cloud based rendering in the client. Further, the paper presents a user study with 54 people evaluating a three-people communication use case and a technical analysis to move towards 3D representations of users. This demo consists of two shared virtual environments to communicate and interact with others, i.e. i) a 360-degree virtual space with users being represented as 2D video streams (with the background removed) and ii) a 3D space with users being represented as point clouds (based on color and depth video data).
Simon Gunkel, Hans Stokking, Martin Prins, Nanda van der Stap, Frank Bart ter Haar, Omar Niamut
MMSys6
2018 Towards Low-Complexity Scalable Coding for Ultra-High Resolution Video And Beyond
abstract
The current state of the art of video coding lacks specific tools to efficiently deal with ultra-high resolutions such as 8K video both at a hardware and software level. At the same time, the increase in video resolutions combined with a larger variety in display resolutions results in a need for spatial scalability at low computational complexity. The proposed solution allows for legacy 4K encoders and decoders both software and in particular hardware to be still exploited with 8K content. The proposed format also offers spatial scalability for an average BD-rate increase of 16%. This scalability is achieved by a polyphase subsampling of the input sequence and by leveraging the existing temporal scalability to induce spatial scalability.
Emmanuel Thomas, Alexandre Gabriel, Omar Niamut, Sylvie Dijkstra-Soudarissanane
VCIP3
2017 AltMM 2017 - 2nd International Workshop on Multimedia Alternate Realities
abstract
AltMM 2017, the 2nd International Workshop on Multimedia Alternate Realities at ACM Multimedia aims to provide a forum for researchers and practitioners concerned with multimedia that enables experiencing "alternate realities". Such experiences may allow us to access other worlds, to live other people's stories, to communicate with or experience alternate realities. Different spaces, times or situations can be entered thanks to multimedia contents and systems, which coexist with our current reality, and are sometimes so vivid and engaging that we feel we are living in them. Advances in multimedia are making it possible to create immersive experiences that may involve the user in a different or augmented world, as an alternate reality.
Teresa Chambel, Rene Kaiser, Omar Niamut, Wei Tsang Ooi
ACM Multimedia3
2017 WebVR meets WebRTC: Towards 360-degree social VR experiences
abstract
Virtual Reality (VR) and 360-degree video are reshaping the media landscape, creating a fertile business environment. During 2016 new 360-degree cameras and VR headsets entered the consumer market, distribution platforms are being established and new production studios are emerging. VR is evermore becoming a hot topic in research and industry and many new and exciting interactive VR content and experiences are emerging. The biggest gap we see in these experiences are social and shared aspects of VR. In this demo we present our ongoing efforts towards social and shared VR by developing a modular web based VR framework, that extends current video conferencing capabilities with new functionalities of Virtual and Mixed Reality. It allows us to connect two people together for mediated audio-visual interaction, while being able to engage in interactive content. Our framework allows to run extensive technological and user based trials in order to evaluate VR experiences and to build immersive multi-user interaction spaces. Our first results indicate that a high level of engagement and interaction between users is possible in our 360-degree VR set-up utilizing current web technologies.
Simon Gunkel, Martin Prins, Hans Stokking, Omar Niamut
VR4
2016 AltMM 2016: 1st International Workshop on Multimedia Alternate Realities
abstract
Multimedia experiences allow us to access other worlds, to live other people's stories, to communicate with or experience alternate realities. Different spaces, times or situations can be entered thanks to multimedia contents and systems, which coexist with our current reality, and are sometimes so vivid and engaging that we feel we are living in them. Advances in multimedia are making it possible to create immersive experiences that may involve the user in a different or augmented world, as an alternate reality. AltMM 2016, the 1st International Workshop on Multimedia Alternate Realities at ACM Multimedia, aims at exploring how the synergy between multimedia technologies and effects can foster the creation of alternate realities and make their access an enriching, valuable and real experience. The workshop program will contain a combination of oral and invited keynote presentations, and poster, demo and discussion sessions, altogether enabling interactive scientific sharing and discussion between practitioners and researchers.
Teresa Chambel, Rene Kaiser, Omar Niamut, Wei Tsang Ooi, Judith Redi
ACM Multimedia3
2016 Using MPEG DASH SRD for zoomable and navigable video
abstract
This paper presents a video streaming client implementation that makes use of the Spatial Relationship Description (SRD) feature of the MPEG-DASH standard, to provide a zoomable and navigable video to an end user. SRD allows a video streaming client to request spatial subparts of a particular video stream, which might be available in multiple resolutions.
Lucia D'Acunto, Jorrit van den Berg, Emmanuel Thomas, Omar Niamut
MMSys4
2016 MPEG DASH SRD: spatial relationship description
abstract
This paper presents the Spatial Representation Description (SRD) feature of the second amendment of MPEG DASH standard part 1, 23009-1:2014 [1]. SRD is an approach for streaming only spatial sub-parts of a video to display devices, in combination with the form of adaptive multi-rate streaming that is intrinsically supported by MPEG DASH. The SRD feature extends the Media Presentation Description (MPD) of MPEG DASH by describing spatial relationships between associated pieces of video content. This enables the DASH client to select and retrieve only those video streams at those resolutions that are relevant to the user experience. The paper describes the design principles behind SRD, the different possibilities it enables and examples of how SRD was used in different experiments on interactive streaming of ultra-high resolution video.
Omar Niamut, Emmanuel Thomas, Lucia D'Acunto, Cyril Concolato, Franck Denoual, Seong Yong Lim
MMSys1
2013 Towards a format-agnostic approach for production, delivery and rendering of immersive media
abstract
The media industry is currently being pulled in the often-opposing directions of increased realism (high resolution, stereoscopic, large screen) and personalization (selection and control of content, availability on many devices). We investigate the feasibility of an end-to-end format-agnostic approach to support both these trends. In this paper, different aspects of a format-agnostic capture, production, delivery and rendering system are discussed. At the capture stage, the concept of layered scene representation is introduced, including panoramic video and 3D audio capture. At the analysis stage, a virtual director component is discussed that allows for automatic execution of cinematographic principles, using feature tracking and saliency detection. At the delivery stage, resolution-independent audiovisual transport mechanisms for both managed and unmanaged networks are treated. In the rendering stage, a rendering process that includes the manipulation of audiovisual content to match the connected display and loudspeaker properties is introduced. Different parts of the complete system are revisited demonstrating the requirements and the potential of this advanced concept.
Omar Niamut, Axel Kochale, Javier Ruiz Hidalgo, Rene Kaiser, Jens Spille, Jean-François Macq, Gert Kienast, Oliver Schreer, Ben Shirley
MMSys1
2006 RD Optimal Temporal Noise Shaping for Transform Audio Coding
abstract
In this article we investigate rate-distortion optimal temporal noise shaping for transform audio coding. Temporal noise shaping, or TNS, is a technique for reshaping the quantization noise in the time domain through open-loop linear predictive coding of frequency domain coefficients. Traditionally, a selection mechanism based on prediction gain is employed to determine whether it is advantageous to apply TNS or not. Although this method is effective for reducing coding artifacts in transient and speech signals, critical adjustment of the prediction gain threshold is necessary to avoid excessive bit rate demands. We propose the use of TNS in a rate-distortion optimization framework. Within this framework a jointly optimal selection of the prediction filter order and the quantizer for coding the coefficients can be made, such that the perceptual distortion is minimized for a given target rate. Experimental results for an MDCT-based audio coding system are presented and it is shown that TNS within an RD optimization framework outperforms the existing TNS method
Omar Niamut, Richard Heusdens
ICASSP (5)1
2006 Perceptual Audio Coding Using N-Channel Lattice Vector Quantization
abstract
We consider the problem of reliable distribution of audio over packet-switched networks. We make use of multiple-description coding combined with transform coding in order to obtain robustness towards packet losses. Previous approaches to this problem were restricted to the case of only two descriptions. In this work we use n-channel multiple-description lattice vector quantizers (MD-LVQs), which allow for the possibility of using more than two descriptions. For a given packet-loss probability we find the number of descriptions and the bit allocation between transform coefficients which minimizes a perceptual distortion measure subject to an entropy constraint. The optimal quantizers are presented in closed form, thus avoiding any iterative quantizer design procedures. The theoretical results are verified with numerical computer simulations using audio signals and it is shown that in environments with excessive packet losses it is advantageous to use more than two descriptions. We verify in subjective listening tests that using more than two descriptions lead to signals of perceptually higher quality
Jan Østergaard, Omar Niamut, Jesper Jensen 0001, Richard Heusdens
ICASSP (5)2
2005 Optimal time segmentation for overlap-add systems with variable amount of window overlap
abstract
In this letter, we propose a new best basis search algorithm for computing the optimal time segmentation of a signal, given a predefined cost measure. The new algorithm solves a problem that arises when the individual signal segments are windowed and overlap-add is applied between adjacent signal segments. When windows having a variable tail shape are employed, the minimization of a cost measure is faced with dependencies between segmental costs due to varying window overlap. A dynamic programming-based algorithm is presented that takes into account these dependencies. It computes both the optimal split positions and the optimal amount of window overlap at these split positions in polynomial time. The proposed algorithm gives an upper bound to the achievable performance of existing algorithms. Experimental results for a modified discrete cosine transform-based processing system are presented, both for entropy and rate-distortion cost measures. These results show a performance gain over existing schemes at the cost of an increased computational complexity.
Omar Niamut, Richard Heusdens
IEEE Signal Process. Lett.1
2003 Flexible frequency decompositions for cosine-modulated filter banks
abstract
We investigate the use of nonuniform cosine-modulated filter banks for audio coding. A rate-distortion framework is employed, similar to the work in Herley et al. (1994), to select the filter bank structure from a large library of possible frequency decompositions. A new flexible frequency decomposition algorithm is proposed that jointly optimizes the filter bank structure and the bit allocation over the subband channels. Experimental results for both synthetic and real audio signals are provided. The new algorithm shows significant improvements in comparison with fixed uniform frequency decompositions, but special care has to be taken to reduce the size of the decomposition overhead.
Omar Niamut, Richard Heusdens
ICASSP (5)1
2003 Subband merging in cosine-modulated filter banks
abstract
Recently, a new method for constructing nonuniform modulated lapped transforms (MLTs) was introduced, by combining subband filters of a uniform MLT. The design, however, was restricted to combining two or four subband filters only, and no systematic design procedure was given. In this letter, we propose an extension to the above-mentioned method that allows arbitrary numbers of subbands to be combined in a systematic way. We investigate the general case of combining filters in arbitrary cosine-modulated filter banks, and give conditions on how to combine the constituent filters such that the resulting nonuniform filter banks have suitable frequency responses.
Omar Niamut, Richard Heusdens
IEEE Signal Process. Lett.1