Sergi Fernández

dblp:18/10714 · also Sergi Fernández Langa · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-9138-0875ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Computer networks · 2
YearPublicationVenuePosition
2026 Towards Interactive Volumetric Video experiences on Lightweight VR Devices through Remote Rendering
abstract
Immersive Virtual Reality (VR) and Volumetric Video (VV) technologies enable highly realistic and interactive experiences but impose significant computational and bandwidth demands. This paper presents a comprehensive end-to-end architecture, along associated technological components and workflows specifically designed to support these advanced services and experiences on lightweight devices like smartphones and even standalone VR headsets via Remote Rendering. By offloading demanding processing tasks to a Remote Renderer, the proposed architecture is able to interactively provide low-latency rendered streams from complex VR scenes including live VV feeds (e.g., holoported users), enabling seamless participation via browser-based thin clients. The proposed framework includes the exploration of advanced stereoscopic delivery strategies, supporting viewport-aware projections to concentrate visual quality within the user’s Field of View (FoV) while enabling seamless exploration with reduced bandwidth overhead. Furthermore, it incorporates 6 Degrees of Freedom (6DoF) capabilities and diverse interaction features from the thin clients. Preliminary evaluations demonstrate promising latency outcomes, achieving end-to-end latency levels under 40ms in local scenarios and motion-to-photon latency levels around 180 ms over the Internet, as well as system stability regardless of VR scene complexity. Ultimately, this work proposes a robust foundation for democratizing premium VR experiences, actively supporting versatile one-to-many delivery scenarios.
Miguel Fernández-Dasí, Antonio Calvo García del Valle, Ernesto Fontes, Sergi Fernández, Josep Paradells Aspas, Mario Montagud
IMX4
2025 Experimental Assessment of Neural 3D Reconstruction for Small UAV-based Applications
abstract
The increasing miniaturization of Unmanned Aerial Vehicles (UAVs) has expanded their deployment potential to indoor and hard-to-reach areas. However, this trend introduces distinct challenges, particularly in terms of flight dynamics and power consumption, which limit the UAVs’ autonomy and mission capabilities. This paper presents a novel approach to overcoming these limitations by integrating Neural 3D Reconstruction (N3DR) with small UAV systems for fine-grained 3-Dimensional (3D) digital reconstruction of small static objects. Specifically, we design, implement, and evaluate an N3DR-based pipeline that leverages advanced models, i.e., Instant-ngp, Nerfacto, and Splatfacto, to improve the quality of 3D reconstructions using images of the object captured by a fleet of small UAVs. We assess the performance of the considered models using various imagery and pointcloud metrics, comparing them against the baseline Structure from Motion (SfM) algorithm. The experimental results demonstrate that the N3DR-enhanced pipeline significantly improves reconstruction quality, making it feasible for small UAVs to support high-precision 3D mapping and anomaly detection in constrained environments. In more general terms, our results highlight the potential of N3DR in advancing the capabilities of miniaturized UAV systems.
Genís Castillo Gómez-Raya, Álmos Veres-Vitályos, Filip Lemic, Pablo Royo, Mario Montagud, Sergi Fernández, Sergi Abadal, Xavier Pérez Costa
PIMRC6
2025 Perceptual Quality Assessment of Compressed Volumetric Human Holographic Sequences at Varying Viewing Distances in VR
abstract
Volumetric video compression is essential for addressing processing and bandwidth challenges, while minimizing perceptual quality loss, in emerging holographic tele-transportation (i.e. holo-portation) services. Research in this field has primarily focused on compression efficiency and/or on real-time performance, but less attention has been given to the perceptual impact of volumetric video compression when viewing dynamic sequences of human holograms at varying distances in Virtual Reality (VR) scenarios. This paper first provides objective evidence on the promising performance of a novel Video-based Point Cloud Compression Codec (V*-PCC) compared to a state-of-the-art Geometry / anchor octree-based (G-PCC) codec. Then, the V*-PCC codec is adopted to run a user study in which observers view sequences of volumetric holographic user representations at different compression levels, shown at varying distances. The results reveal two key insights: (i) perceptual quality improves at equivalent compression settings as the viewing distance increases; and (ii) compression artifacts become less noticeable and more tolerable at greater distances. These findings not only support the need for further Quality of Experience (QoE) studies but interestingly reinforce the potential of devising dynamic viewport- and position-aware encoding strategies to overcome current scalability and interoperability limitations in this field.
Mohamad Hjeij, Mario Montagud, David Rincón Rivera, Leonel Toledo, Sergi Fernández
QoMEX5
2025 Social eXtended Reality (XR) and Virtual Production: Toward New Engaging Immersive Experiences
abstract
Social eXtended Reality (XR) is poised to become a dominant medium for remote communication, social interaction, and collaboration in the near future.However, its potential can be significantly magnified by integrating additional technological enablers.On the one hand, the support for realistic and volumetric holographic representations of users, captured in real-time via affordable setups, will result in enhanced levels of quality of interaction, co-presence and trustworthiness compared to using synthetic avatar-based user representation formats.On the other hand, the availability of highquality multi-modal virtual production tools will enrich the immersion, interaction and storytelling possibilities, while it will allow providing such engaging services to large-scale audiences via 2D video distribution platforms.This paper presents a strategic vision toward achieving a modular and seamless integration between innovative Social XR, holographic communication, and virtual production tools to exploit all such potential advantages.In particular, it proposes a high-level architectural framework to cohesively integrate such technological enablers, accommodating novel features to meet newly derived requirements.Then, it outlines potential applicability scenarios and reports preliminary results as initial evidence of the full achievable potential.
Mario Montagud, Álvaro Egea Benavente, Marc Martos Cabré, Javier Montesa, Francisco Ibañez, Sergi Fernández
IMX6
2024 Content format and quality of experience in virtual reality
Henrique Debarba, Mario Montagud, Sylvain Chagué, Javier Garcia-Lajara Herrero, Ignacio Lacosta, Sergi Fernández, Caecilia Charbonnier
Multim. Tools Appl.6
2023 Addressing Scalability for Real-time Multiuser Holo-portation: Introducing and Assessing a Multipoint Control Unit (MCU) for Volumetric Video
abstract
Scalability, interoperability, and cost efficiency are key remaining challenges to successfully providing real-time holo-portation (and Metaverse-like) services. This paper, for the first time, presents the design and integration of a Multipoint Control Unit (MCU) in a pioneering real-time holo-portation platform, supporting realistic and volumetric user representations (i.e., 3D holograms), with the aim of overcoming such challenges. The feasibility and implications of adopting such an MCU, in comparison with state-of-the-art architectural approaches, are assessed through experimentation in two different deployment setups, by iteratively increasing the number of concurrent users in shared sessions. The obtained results are promising, as it is empirically proved that the newly adopted stream multiplexing together with the novel per-client and per-frame Volumetric Video (VV) processing optimization features provided by the MCU allow increasing the number of concurrent users, while: (i) significantly reducing resources consumption metrics (e.g., CPU, GPU, bandwidth) and frame rate degradation on the client side; and (ii) keeping the end-to-end latency within acceptable limits.
Sergi Fernández, Mario Montagud, David Rincón Rivera, Juame Moragues, Gianluca Cernigliaro
ACM Multimedia1
2023 Enabling and Understanding Interactive Social VR360 Video Viewing
abstract
This paper reports on the research being done towards enabling and understanding interactive social VR360 video viewing scenarios, by exclusively relying on web-based technologies, and using different types of consumption devices. After motivating the relevance of the research topic and associated impact, the paper elaborates on key requirements, features, and system components to effectively enable such scenarios, such as: adaptive and low-latency streaming, media synchronization, social presence, interaction channels, and assistive methods. For each of these features and components, different alternatives are assessed and proof of concept implementations are being provided. With an effective combination and integration of all these contributions, an end-to-end platform can be built and used as a research framework to explore the applicability and potential benefits of social VR360viewing in a variety of use cases, like education, culture or surveillance, by tailoring the technological components based on lessons learned from experimental studies. These use case studies can also provide relevant insights into activity patterns, behaviors, and preferences in Social Viewing scenarios.
Miguel Fernández-Dasí, Mario Montagud, Isaac Fraile, Josep Paradells Aspas, Sergi Fernández
IMX5
2020 PC-MCU: point cloud multipoint control unit for multi-user holoconferencing systems
abstract
This paper introduces the Point Cloud Multipoint Control Unit (PC-MCU): a key component for multi-user holoconferencing systems, where remote participants are represented as Point Clouds. The presented solution redefines the idea of MCU, broadly used to optimize connections and communications between users in traditional videoconferencing, and introduces a set of key features for the optimization of holoconferencing services where multiple users can be remotely connected. The PC-MCU is a virtualized cloud-based component, that aims at reducing the end-user client computational resources and bandwidth usage, providing the following key features: fusion of volumetric videos, Level of Detail (LoD) adjustment and non visible data removal. The results obtained for a scenario with two remote users, show how the introduction of the PC-MCU provides significant benefits in terms of computational resources and bandwidth savings, thus alleviating the requirements at the client side in holoconferencing services when compared to a baseline condition without using it. These improvements open the door to further research on this area to enable scalable and adaptive holoconferencing services using lightweight devices.
Gianluca Cernigliaro, Marc Martos Cabré, Mario Montagud, Amir Ansari, Sergi Fernández
NOSSDAV5
2018 ImmersiaTV: enabling customizable and immersive multi-screen TV experiences
abstract
ImmersiaTV is a H2020 European project that targets the creation of novel forms of TV content production, delivery and consumption to enable customizable and immersive multi-screen TV experiences. The goal is not only to provide an efficient support for multi-screen scenarios, but also to achieve a seamless integration between the traditional TV content formats and consumption devices with the emerging omnidirectional ones, thus opening the door to new fascinating scenarios. This paper initially provides an overview of the end-to-end platform that is being developed in the project. Then, the created contents and considered pilot scenarios are briefly described. Finally, the paper provides details about the consumption part of the ImmersiaTV platform to be showcased. In particular, it enables a customizable, interactive and synchronized consumption of traditional and omnidirectional contents from an opera performance, in multiscreen scenarios, composed of main TVs, tablets and Head Mounted Displays (HMDs).
David Gómez 0003, Juan A. Núñez, Mario Montagud, Sergi Fernández
MMSys4
2018 TiCMP: A lightweight and efficient Tiled Cubemap projection strategy for Immersive Videos in Web-based players
abstract
The encoding, delivery and interactive consumption of omnidirectional videos still face many challenges. Traditional encoding techniques, based on Equirectangular Projection (ERP) formats, introduce significant pixel redundancy. This has prompted the appearance of advanced solutions based on the segmentation in separate regions or tiles, and their selective delivery depending on the users' viewpoint. However, tiling techniques introduce further challenges. First, neighboring pixels are encoded separately, which may result in noticeable separations between regions. Second, they can involve synchronization problems when the users' viewpoints change. Third, they may require further extensions to existing technologies, such as Dynamic Adaptive Streaming over HTTP (DASH), which makes their adoption in current web browsers very challenging. The use of Cubemap Projection (CMP) is alternatively gaining popularity due to its advantages compared to ERP. However, it requires the streaming of the whole 360° area. This paper proposes a novel tiled Cubemap (TiCMP) strategy that overcomes all the mentioned limitations. TiCMP is based on dividing the cube into two tiles, adaptively streaming them based on the users' viewpoint, and playing them out in a synchronized manner in web-based players. Evaluation results demonstrate that TiCMP provides significant bandwidth savings, without negatively impacting the Quality of Experience (QoE) when compared to traditional Equirectangular- and Cubemap-based strategies.
David Gómez 0003, Juan A. Núñez, Isaac Fraile, Mario Montagud, Sergi Fernández
NOSSDAV5