VLDB 2026 Research / reviewers in the wild / expert
Pablo César
dblp:03/6828
· DBLP profile ↗
139ranked-venue papers
19as first author
58since 2021 · last 2026
0000-0003-1752-6837ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 94 · 14 first-author · 43 since 2021Human-computer interaction and ubiquitous computing · 44 · 27 since 2021Computer networks · 17 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 13 · 4 first-authorArtificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | More Human or More AI? Visualizing Human-AI Collaboration Disclosures in Journalistic News ProductionabstractWithin journalistic editorial processes, disclosing AI usage is currently limited to simplistic labels, which misses the nuance of how humans and AI collaborated on a news article. Through co-design sessions (N=10), we elicited 69 disclosure designs and implemented four prototypes that visually disclose human–AI collaboration in journalism. We then ran a within-subjects lab study (N=32) to examine how disclosure visualizations (Textual, Role-based Timeline, Task-based Timeline, Chatbot) and collaboration ratios (Primarily Human vs. Primarily AI) influenced visualization perceptions, gaze patterns, and post-experience responses. We found that textual disclosures were least effective in communicating human-AI collaboration, whereas Chatbot offered the most in-depth information. Furthermore, while role-based timelines amplified AI contribution in primarily human articles, task-based timeline shifted perceptions toward human involvement in primarily AI articles. We contribute Human-AI collaboration disclosure visualizations and their evaluation, and cautionary considerations on how visualizations can alter perceptions of AI’s actual role during news article creation. Amber Kusters, Pooja Prajod, Pablo César, Abdallah El Ali |
CHI | 3 |
| 2026 | WebRTC-Based Volumetric Video Conferencing: SFU Architecture Evaluation and BenchmarkingabstractImmersive technologies promise to revolutionize communication through enhanced sense of presence and interactivity. To enable interaction, reliable low-latency transport mechanisms are needed to handle the large volumes of data created by complex 3D objects. In this paper, we propose an open-source, codec-independent, selective forwarding unit (SFU) for real-time volumetric video streaming using WebRTC. For evaluation purposes, we provide a reference client implementation by extending VR2Gather, a TCP-based system for immersive communication. We conduct extensive evaluations using both new and existing datasets to compare the performance of WebRTC against TCP-based protocols in an emulated testbed environment. The evaluations demonstrate that WebRTC outperforms other protocols in high-latency scenarios and adapts video quality to user movement 13% and 36% faster than its TCP-based counterparts in networks with 5 ms and 10 ms of network latency, respectively. Matthias De Fré, Jeroen van der Hooft, Jack Jansen 0001, Silvia Rossi 0001, Thomas Röggla, Tim Wauters, Filip De Turck, Irene Viola 0001, Pablo César |
NOSSDAV | 9 |
| 2026 | Social XR for Pre-Production Meetings: An In-the-Wild Study of Early Stage Communication Between XR Producers and Clients
Sueyoon Lee, Irene Viola 0001, Jack Jansen 0001, Ashutosh Singla, Karolina Wylezek, Thomas Röggla, Pablo César |
IMX | 7 |
| 2026 | A Comparative Study of VR and Tablet Interactions for an Immersive Jazz Concert with Spatial Audio
Ibrahim El Shemy, Patricia de Torres Coll, Katrien De Moor, Francesco Mariotti, Anna Bindi, Rob Oldfield, Karolina Wylezek, Pablo César |
IMX | 8 |
| 2026 | The Influence of Context on Learning in a Social VR Historical Fashion Exhibition
Karolina Wylezek, Irene Viola 0001, Silvia Rossi 0001, Jack Jansen 0001, Thomas Röggla, Pablo César |
IMX | 6 |
| 2026 | Understanding trust toward human versus AI-generated health information through behavioral and physiological sensingabstractAs AI-generated health information proliferates online and becomes increasingly indistinguishable from human-sourced information, it becomes critical to understand how people trust and label such content, especially when the information is inaccurate. We conducted two complementary studies: (1) a mixed-methods survey (N=142) employing a 2 (source: Human vs. LLM) × 2 (label: Human vs. AI) × 3 (type: General, Symptom, Treatment) design, and (2) a within-subjects lab study (N=40) incorporating eye-tracking and physiological sensing (ECG, EDA, skin temperature). Participants were presented with health information varying by source-label combinations and asked to rate their trust, while their gaze behavior and physiological signals were recorded. We found that LLM-generated information was trusted more than human-generated content, whereas information labeled as human was trusted more than that labeled as AI. Trust remained consistent across information types. Eye-tracking and physiological responses varied significantly by source and label. Machine learning models trained on these behavioral and physiological features predicted binary self-reported trust levels with 73 % accuracy and information source with 65 % accuracy. Our findings demonstrate that adding transparency labels to online health information modulates trust. Behavioral and physiological features show potential to verify trust perceptions and indicate if additional transparency is needed. Xin Sun 0016, Rongjun Ma, Shu Wei, Pablo César, Jos A. Bosch, Abdallah El Ali |
Int. J. Hum. Comput. Stud. | 4 |
| 2025 | Haptic Biosignals Affect Proxemics Toward Virtual Reality AgentsabstractEncounters with virtual agents currently lack the haptic viscerality of human contact. While digital biosignal communication can medi-ate such virtual social interactions, how artifcial haptic biosignals infuence users personal space during Virtual Reality (VR) experi-ences is unknown. Designing vibrotactile heartbeats and thermally-actuated body temperature, we ran a within-subjects study (N=31) to investigate feedback (Thermal, Vibration, Thermal+Vibration, None) and agent stories (Negative, Neutral, Positive) on objective and subjective interpersonal distance (IPD), perceived arousal and comfort, presence, and post-experience responses. Findings showed that thermal feedback decreased objective but not subjective IPD, whereas vibrotactile heartbeats (signaling agent's closeness) increased both while heightening arousal and discomfort. Agents stories did not afect IPD, arousal, or comfort. Our qualitative fndings shed light on signal ambiguity and presence constructs within VR-based haptic stimulation. We contribute insights into artifcial biosignals and their infuence on VR proxemics, with cautionary considerations should the boundaries blur between physical and virtual touch. Simone Ooms, Minha Lee, Ekaterina R. Stepanova, Pablo César, Abdallah El Ali |
CHI | 4 |
| 2025 | QoE Evaluation of Remote Physiotherapy in Volumetric Video and Video-Based Real-Time CommunicationabstractIn recent years, video conferencing platforms have become powerful tools for remote communication. There has also been an increase in the use of VR systems for communication. However, very few of these systems utilize photorealistic human representation. This paper investigates the strengths, challenges, and limitations of a novel 3D communication prototype (VR2Gather) and a well-established video conferencing system (Zoom). Specifically, we explore whether the 3D communication prototype can achieve comparable performance levels in a remote physiotherapy use case. By assessing various aspects, such as audio-visual quality, presence, and interaction, we aim to determine if the current prototype is comparable with commercial systems in some dimensions while exceeding expectations in others. Our results indicated that VR2Gather has the potential for a better sense of connection and higher concentration. However, challenges like improving 3D rendering quality and communication ease still need to be overcome to make it suitable for physiotherapy. Ashutosh Singla, Irene Viola 0001, Jack Jansen 0001, Pablo César |
ICME | 4 |
| 2025 | UVG-CWI-DQPC: Dual-Quality Point Cloud Dataset for Volumetric Video ApplicationsabstractVolumetric video is a key enabler of immersive extended reality (XR) experiences and is often represented using point clouds for their structural simplicity. However, capturing volumetric content through multi-view acquisition and depth sensing poses many challenges, such as occlusions and depth mismatches. To foster research in this field, we introduce a unique dual-quality point cloud dataset, named UVG-CWI-DQPC, which is designed to support the development of point cloud enhancement, compression, and quality assessment. Our dataset includes 12 dynamic sequences captured simultaneously by: 1) a high-end capture system producing high-fidelity point clouds with extensive processing; and 2) a consumer-grade capture system relying on affordable RGB-D cameras, lightweight processing, and open-source tools. For each sequence, our dataset provides ground-truth point clouds from the high-end capture system and raw RGB-D footage from the consumer-grade capture system, along with calibration data and tools for point cloud generation. This dual-quality setup enables direct comparison and benchmarking of algorithms for densification, occlusion removal, registration, and quality enhancement. Our dataset is publicly available under a permissive license to support reproducible research and standardization work in Moving Picture Experts Group (MPEG) and 3rd Generation Partnership Project (3GPP). Guillaume Gautier, Xuemei Zhou, Jack Jansen 0001, Louis Fréneau, Marko Viitanen, Uyen Phan, Jani Käpylä, Irene Viola 0001, Alexandre Mercat, Pablo César, Jarno Vanne |
ACM Multimedia | 11 |
| 2025 | RCQoEA-360VR: Real-time Continuous QoE Scores for HMD-based 360° VR DatasetabstractAs immersive 360° video experiences through head-mounted displays (HMDs) gain widespread adoption, the need for real-time, fine-grained assessment of Quality of Experience (QoE) becomes increasingly critical for optimising user engagement and system performance. This paper introduces RCQoEA-360VR, a novel multi-modal dataset designed for continuous QoE evaluation in virtual reality (VR) environments. In a controlled study (N=32), participants watched five selected 360° video sequences across eight different video quality configurations (from the VQEG database) using a Vive Pro Eye while providing continuous QoE annotations via a touchpad-based input method, enhanced by the DotMorph peripheral visualisation technique. The dataset also includes synchronised physiological signals (electrocardiogram and galvanic skin response), behavioural data (eye and head movements) and post-viewing QoE ratings gathered through a within-VR interface. RCQoEA-360VR addresses a critical gap in existing public datasets by providing a fine-grained, synchronised multimodal data for immersive QoE analysis. It offers a unique and valuable resource for the research community, supporting a wide range of research applications, including QoE prediction, behavioural modelling, adaptive streaming, and implicit perceptual analysis. Sowmya Vijayakumar, Tong Xue, Abdallah El Ali, Irene Viola 0001, Ronan Flynn, Peter Corcoran 0001, Pablo César, Niall Murray |
ACM Multimedia | 7 |
| 2025 | Understanding AI Disclosure Needs for News Production and JournalismabstractArtificial Intelligence (AI) is revolutionizing the way content is produced and integrated into journalistic workflows. The EU AI act’s Article 50 sets up transparency requirements aimed at encouraging the adoption and disclosure of AI in an ethical and responsible manner. In this study, we organized focus group interviews with Dutch citizens (N=21) to understand their expectations and needs regarding AI disclosures in the context of news production and journalism. These conversations are essential to understand if legal and regulatory policies are grounded in real-world experiences of citizens, and adequately address their concerns and enhance their digital interactions. We found that citizens predominantly favor disclosures of AI usage in journalistic content, in the form of (1) source references, (2) visual indicators (logos/watermarks) and (3) have varying preferences regarding information presentation and interaction modalities. Our findings highlight the need for interdisciplinary approaches to align standardization efforts with AI disclosures for news media. Karthikeya Puttur Venkatraj, Sophie Morosoli, Hannes Cools, Laurens Naudts, Claes H. de Vreese, Natali Helberger, Pablo César, Abdallah El Ali |
MUM | 7 |
| 2025 | From Individual QoE to Shared Mental Models: A Novel Evaluation Paradigm for Collaborative XRabstractExtended Reality (XR) systems are rapidly shifting from isolated, single-user applications towards collaborative and social multi-user experiences. To evaluate the quality and effectiveness of such interactions, it is therefore required to move beyond traditional individual metrics such as Quality-of-Experience (QoE) or Sense of Presence (SoP). Instead, group-level dynamics such as effective communication, coordination etc. need to be encompassed to assess the shared understanding of goals and procedures. In psychology, this is referred to as a Shared Mental Model (SMM). The strength and congruence of such an SMM are known to be key for effective team collaboration and performance. In an immersive XR setting, though, novel Influence Factors (IFs) emerge that are not considered in a setting of physical co-location. Evaluations on the impact of these novel factors on SMM formation in XR, however, are close to non-existent. Therefore, this work proposes SMMs as a novel evaluation tool for collaborative and social XR experiences. To better understand how to explore this construct, we ran a prototypical experiment based on ITU recommendations in which the influence of asymmetric end-to-end latency is evaluated through a collaborative, two-user block building task. The results show how also in an XR context strong SMM formation can take place even when collaborators have fundamentally different responsibilities and behavior. Moreover, the study confirms previous findings by showing in an XR context that a teams’ SMM strength is positively associated with its performance. Sam Van Damme, Jack Jansen 0001, Silvia Rossi 0001, Pablo César |
QoMEX | 4 |
| 2025 | PhysioDrum: Bridging Physical and Digital Realms in Immersive Musical InteractionabstractThe Internet of Multisensory, Multimedia, and Musical Things (Io3MT) bridges computer science, humanities, and arts, fostering transmedia services and creative applications.This demo research applies these principles alongside extended reality (XR) to enhance PhysioDrum, an immersive, multimodal system that blends physical and digital aspects to expand musical expression in virtual environments.Using a smart musical instrument (SMI) and electronic pedals as interfaces, users interact with a virtual drum kit through gestures while receiving haptic feedback.By integrating sound and multimedia elements, PhysioDrum aims to reduces cognitive load and the learning curve, merging traditional drumming practices with immersive XR.The demo emphasizes design strategies that enhance playability, accessibility, and creative potential for users of all skill levels. Rômulo Vieira, Débora C. Muchaluat-Saade, Pablo César |
IMX | 3 |
| 2025 | Enhancing the Audience Experience for VR and AR Theatre with AI-generated SubtitlesabstractRecent technological developments on AI and immersive media are transforming the artistic landscape, providing novel mechanisms for artists and audiences. Following a human-centric approach, together with a theatre company in Greece, this paper investigates how subtitle placement affects user experience and cognitive load in a live theatre performance enhanced by AR glasses. To do so, we design and develop a system for displaying subtitles in VR and AR. We evaluated the system in two conditions (N = 19;N = 12), both in a controlled environment (VR) and an actual theatre (AR). In the latter, we integrate AI solutions to provide automatic captioning and translation in real time, and VFX to further augment the experience. Our quantitative and qualitative results showed no difference between subtitle placements in terms of cognitive load and user experience, with users equally liking the two proposed approaches. Results also highlighted the perceived usefulness of AR to enhance theatre performances, indicating new paths for wider accessibility and further immersion. Irene Viola 0001, Moonisa Ahsan, Olga Chatzifoti, Atanas Yonkov, Eleni Oikonomou, Ioannis Radin, Pawel Maka, Abderrahmane Issam, Pablo César |
VRST | 9 |
| 2025 | VRD: A multi-lingual translation and Ai-Assisted Navigation experience for VR Conference ApplicationabstractThis paper presents an AI-assisted VR conference application with multilingual translation and navigation agent capabilities. A pilot study with 18 participants (11 females, 7 males) was conducted to assess the system’s usability. AI-assisted navigation worked smoothly, but the AI translation had issues that prevented the users from having a good experience, nonetheless, participants expressed positive attitudes toward the system, and future work will focus on achieving better user experience. Moonisa Ahsan, Irene Viola 0001, Manuel Toledo, Dimitris Kontopoulos, Pablo César |
VRST | 5 |
| 2025 | PointPCA+: A full-reference Point Cloud Quality Assessment metric with PCA-based features
Xuemei Zhou, Evangelos Alexiou, Irene Viola 0001, Pablo César |
Signal Process. Image Commun. | 4 |
| 2025 | Adaptive Cloud VR Gaming Optimized by Gamer QoE ModelsabstractCloud Virtual Reality (VR) gaming offloads computationally intensive VR games to resourceful data centers. However, ensuring good Quality of Experience (QoE) in cloud VR gaming is inherently challenging as VR gamers demand high visual quality, short response time, and negligible cybersickness. In this article, we study the QoE of cloud VR gaming and build a QoE-optimized system in a few steps. First, we establish a cloud VR gaming testbed capable of emulating various network conditions. Using the testbed, we conduct comprehensive QoE evaluations using a user study to evaluate the influence of diverse factors, such as encoding settings, network conditions, and game genres, on gamer QoE scores. Second, we construct the very first QoE models for cloud VR gaming using our QoE evaluation results. Our QoE models achieve up to 0.93 ( \(\sigma=0.02\) ) in Pearson Linear Correlation Coefficient (PLCC) and 0.92 ( \(\sigma=0.02\) ) in Spearman Rank-Order Correlation Coefficient (SROCC), where \(\sigma\) stands for the standard deviation. Last, we leverage our QoE models for dynamically adapting encoding settings in our testbed. Extensive experiments revealed that, compared to the current practice, our adaptive cloud VR gaming system improves: (i) overall quality by 0.87 ( \(\sigma=0.44\) ), (ii) visual quality by 0.61 ( \(\sigma=0.45\) ), and (iii) interaction quality by 1.20 ( \(\sigma=0.48\) ) on average in 5-point Mean Opinion Score (MOS). Kuan-Yu Lee, Ashutosh Singla, Pablo César, Cheng-Hsin Hsu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | A Clustering Approach to Unveil User Similarities in 6 df Extended Reality ApplicationsabstractThe advent in our daily life of Extended Reality (XR) technologies, such as Virtual and Augmented Reality, has led to the rise of user-centric systems, offering higher level of interaction and presence in virtual environments. In this context, understanding the actual interactivity of users is still an open challenge and a key step to enabling user-centric system. In this work, our goal is to construct an efficient clustering tool for 6 df navigation trajectories by extending the applicability of existing behavioural tool. Specifically, we first compare the navigation in 6 df with its 3 df counterpart, highlighting the main differences and novelties. Then, we investigate new metrics aimed at better modelling behavioural similarities between users in a 6 df system. More concretely, we define and compare 11 similarity metrics which are based on different distance features (i.e., user positions in the 3D space, user viewing directions) and distance measurements (i.e., Euclidean, Geodesic, angular distance). Our solutions are validated and tested on real navigation paths of users interacting with dynamic volumetric media in both 6 df Virtual Reality and Augmented Reality conditions. Results show that metrics based on both user position and viewing direction better perform in detecting user similarity while navigating in a 6 df system. Such easy-to-use but robust metrics allow us to answer a fundamental question for user-centric systems: ‘How do we detect if users look at the same content in 6 df?’, opening the gate to new solutions based on users interactivity, such as viewport prediction, live streaming services optimised based on users behaviour but also for user-based quality assessment methods. Silvia Rossi 0001, Irene Viola 0001, Laura Toni, Pablo César |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Subjective and Objective Quality Assessment for Dynamic Point Cloud with Visual Attention in 6 DoFabstractPerceptual quality assessment of Dynamic Point Cloud (DPC) contents plays an important role in various Virtual Reality (VR) applications that involve human beings as the end user. Understanding and modeling perceptual quality assessment is greatly enriched by insights from visual attention. However, incorporating aspects of visual attention in DPC quality models is largely unexplored, as ground-truth visual attention data are scarcely available. Besides, testing methods and procedures for collecting visual attention data are still to be agreed on. This article presents a dataset containing subjective opinion scores and visual attention maps of DPCs, collected in a VR environment using eye-tracking technology. Both the quality score and eye-tracking data were collected during a subjective quality assessment experiment, in which subjects were instructed to watch and rate DPCs at various degradation levels under 6 Degrees of Freedom (DoF) inspection, using a head-mounted display. Qualitative interview analysis was also conducted after the experiment. The dataset consists of 50 DPCs, including 5 reference DPCs, with each reference encoded at 3 distortion levels using 3 different codecs (namely G-PCC, V-PCC, CWI-PCL), amounting to a total of 9 degraded version per reference. Additionally, it incorporates 1,000 gaze trials from 40 participants, yielding a total of 15,000 visual attention maps across all the DPCs. We additionally benchmark objective quality metrics originally designed for static point clouds, evaluating their performance in our dataset using two temporal pooling strategies. Furthermore, we employ the visual attention data that are retrieved during our experiment to evaluate whether the performance of widely used objective quality metrics is improved by considering subjective measurements of visual attention. This dataset establishes a link between quality assessment and visual attention within the context of DPC. Moreover, thematic analysis of the interviews helps uncover user behavior and factors impacting perceptual quality for DPC in 6 DoF. This work deepens our understanding of DPC quality assessment and visual attention, driving progress in the realm of VR experiences and perception. Xuemei Zhou, Irene Viola 0001, Evangelos Alexiou, Jack Jansen 0001, Pablo César |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | Comparison of Visual Saliency for Dynamic Point Clouds: Task-free vs. Task-dependentabstractThis paper presents a Task-Free eye-tracking dataset for Dynamic Point Clouds (TF-DPC) aimed at investigating visual attention. The dataset is composed of eye gaze and head movements collected from 24 participants observing 19 scanned dynamic point clouds in a Virtual Reality (VR) environment with 6 degrees of freedom. We compare the visual saliency maps generated from this dataset with those from a prior task-dependent experiment (focused on quality assessment) to explore how high-level tasks influence human visual attention. To measure the similarity between these visual saliency maps, we apply the well-known Pearson correlation coefficient and an adapted version of the Earth Mover's Distance metric, which takes into account both spatial information and the degrees of saliency. Our experimental results provide both qualitative and quantitative insights, revealing significant differences in visual attention due to task influence. This work enhances our understanding of the visual attention for dynamic point cloud (specifically human figures) in VR from gaze and human movement trajectories, and highlights the impact of task-dependent factors, offering valuable guidance for advancing visual saliency models and improving VR perception. Xuemei Zhou, Irene Viola 0001, Silvia Rossi 0001, Pablo César |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Open-Sourcing VR2Gather: A Collaborative Social VR System for Adaptive Multi-Party Real Time CommunicationabstractSocial Virtual Reality is envisioned to transform how individu- als communicate remotely, offering a sense of immersion and co- presence within a virtual space. Current platforms enabling remote social interactions rely on synthetic user representations. We ad- dress this limitation by enabling realistic human representation through volumetric content capture, encoding and transmission. Specifically, we present an extended version of VR2Gather, now a fully open source Unity package, available at https://github.com/ cwi-dis/VR2Gather-acmmm-oss. Our platform is a customisable system to transmit volumetric content in a multi-party real-time environment, easy to integrate into existing applications. Jack Jansen 0001, Thomas Röggla, Silvia Rossi 0001, Irene Viola 0001, Pablo César |
ACM Multimedia | 5 |
| 2024 | Deciphering Perceptual Quality in Colored Point Cloud: Prioritizing Geometry or Texture Distortion?abstractPoint clouds represent one of the prevalent formats for 3D content. Distortions introduced at various stages in the point cloud processing pipeline affect the visual quality, altering their geometric composition, texture information, or both. Understanding and quantifying the impact of the distortion domain on visual quality is vital to driving rate optimization and guiding post-processing steps to improve the quality of experience. In this paper, we propose a multi-task guided multi-modality no reference metric (M3-Unity), which utilizes 4 types of modalities across attributes and dimensionalities to represent point clouds. An attention mechanism establishes inter/intra associations among 3D/2D patches, which can complement each other, yielding local and global features, to fit the highly nonlinear property of the human vision system. A multi-task decoder involving distortion type classification selects the best association among 4 modalities, aiding the regression task and enabling the in-depth analysis of the interplay between geometrical and textural distortions. Furthermore, our framework design and attention strategy enable us to measure the impact of individual attributes and their combinations, providing insights into how these associations contribute particularly in relation to distortion type. Extensive experimental results on 4 datasets consistently outperform the state-of-the-art metrics by a large margin. The code is available at https://github.com/cwi-dis/ACMMM2024-Oral. Xuemei Zhou, Irene Viola 0001, Yunlu Chen, Jiahuan Pei, Pablo César |
ACM Multimedia | 5 |
| 2024 | Exploring Artificial Intelligence for Advancing Performance Processes and Events in Io3MT
Rômulo Vieira, Débora C. Muchaluat-Saade, Pablo César |
MMM (4) | 3 |
| 2024 | ComPEQ-MR: Compressed Point Cloud Dataset with Eye Tracking and Quality Assessment in Mixed RealityabstractPoint clouds (PCs) have attracted researchers and developers due to their ability to provide immersive experiences with six degrees of freedom (6DoF). However, there are still several open issues in understanding the Quality of Experience (QoE) and visual attention of end users while experiencing 6DoF volumetric videos. First, encoding and decoding point clouds require a significant amount of both time and computational resources. Second, QoE prediction models for dynamic point clouds in 6DoF have not yet been developed due to the lack of visual quality databases. Third, visual attention in 6DoF is hardly explored, which impedes research into more sophisticated approaches for adaptive streaming of dynamic point clouds. In this work, we provide an open-source Compressed Point cloud dataset with Eye-tracking and Quality assessment in Mixed Reality (ComPEQ--MR). The dataset comprises four compressed dynamic point clouds processed by Moving Picture Experts Group (MPEG) reference tools (i.e., VPCC and GPCC), each with 12 distortion levels. We also conducted subjective tests to assess the quality of the compressed point clouds with different levels of distortion. The rating scores are attached to ComPEQ--MR so that they can be used to develop QoE prediction models in the context of MR environments. Additionally, eye-tracking data for visual saliency is included in this dataset, which is necessary to predict where people look when watching 3D videos in MR experiences. We collected opinion scores and eye-tracking data from 41 participants, resulting in 2132 responses and 164 visual attention maps in total. The dataset is available at https://ftp.itec.aau.at/datasets/ComPEQ-MR/. Minh Nguyen 0006, Shivi Vats, Xuemei Zhou, Irene Viola 0001, Pablo César, Christian Timmerer, Hermann Hellwagner |
MMSys | 5 |
| 2024 | Enhancing Immersive Experiences through 3D Point Cloud Analysis: A Novel Framework for Applying 2D Visual Saliency Models to 3D Point CloudsabstractIn the new area of immersive multimedia environments, understanding and manipulating visual attention are crucial for enhancing user experience. This study introduces an innovative framework that extends traditional 2D saliency maps to the analysis of 3D point clouds, a step forward in adapting saliency prediction to more complex and immersive environments. Our framework centers on the orthographic projection of 3D point clouds onto 2D planes, enabling the application of established 2D saliency models to this novel context. We further delve into the evaluation of these models on a 3D point cloud eye-tracking dataset, exploring various projection settings and thresholding techniques to maintain the integrity of saliency information in the transition from 2D to 3D. This research not only bridges a gap in applying visual attention models to 3D data but also offers insights into the optimization of quality of experience in immersive multimedia systems. Marouane Tliba, Xuemei Zhou, Irene Viola 0001, Pablo César, Aladine Chetouani, Giuseppe Valenzise, Frédéric Dufaux |
QoMEX | 4 |
| 2024 | Communication Challenges between Clients and Producers of Immersive Media Applications: can Social XR help?abstractExtended Reality (XR) has emerged as a transformative and immersive technology with versatile applications in content creation and consumption. As XR gains popularity, companies eager to adopt it often possess a surface-level understanding, investing significant resources without effectively addressing the genuine needs of end-users. This study explores the current workflows of XR production companies, and the potential of social XR in mitigating challenges throughout the XR production workflow. We present the outcomes of three respective focus group workshops conducted with three XR production companies and their experts (N=17). The results indicate that at every stage of the production, namely pre-production, production, post-production, and post-release, there are communication challenges between producers and clients, as well as different production and post-production specialists. We discuss various aspects of XR concerning the problem and propose novel opportunities offered by social XR to ameliorate those challenges, improving communication and making development more agile. Sueyoon Lee, Irene Viola 0001, Ashutosh Singla, Pablo César |
IMX | 4 |
| 2024 | Exploring Retrospective Annotation in Long-Videos for Emotion RecognitionabstractEmotion Recognition systems are typically trained to classify a given psychophysiological state into emotion categories. Current platforms for emotion ground-truth collection show limitations for real-world scenarios of long-duration content (e.g., > 10m), namely: 1) Real-time annotation tools are distracting and become exhausting in a longer video; 2) Perform retrospective annotation of the whole content in bulk (providing highly coarse annotations); or 3) Are performed by external experts (depending on the number of annotators and their subjective experience). We explore a novel approach, the EmotiphAI Annotator, that allows undisturbed content visualisation and simplifies the annotation process by using segmentation algorithms that select brief clips for emotional annotation retrospectively. We compare three methods for content segmentation based on physiological data (Electrodermal Activity (EDA), emotion-based), scene (time-based), and random (control) selection. The EmotiphAI Annotator attained a B+ System Usability Scale score and low-average mental workload as per the NASA Task Load Index (40%). The reliability of the self-report was analysed by the inter-rater agreement (STD0.3 to 0.8), where the method based on EDA obtained the overall best performance. Patrícia J. Bota, Pablo César, Ana Fred, Hugo Silva 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | Delay Threshold for Social Interaction in Volumetric eXtended Reality CommunicationabstractImmersive technologies like eXtended Reality (XR) are the next step in videoconferencing. In this context, understanding the effect of delay on communication is crucial. This article presents the first study on the impact of delay on collaborative tasks using a realistic Social XR system. Specifically, we design an experiment and evaluate the impact of end-to-end delays of 300, 600, 900, 1,200, and 1,500 ms on the execution of a standardized task involving the collaboration of two remote users that meet in a virtual space and construct block-based shapes. To measure the impact of the delay in this communication scenario, objective and subjective data were collected. As objective data, we measured the time required to execute the tasks and computed conversational characteristics by analyzing the recorded audio signals. As subjective data, a questionnaire was prepared and completed by every user to evaluate different factors such as overall quality, perception of delay, annoyance using the system, level of presence, cybersickness, and other subjective factors associated with social interaction. The results show a clear influence of the delay on the perceived quality and a significant negative effect as the delay increases. Specifically, the results indicate that the acceptable threshold for end-to-end delay should not exceed 900 ms. This article additionally provides guidelines for developing standardized XR tasks for assessing interaction in Social XR environments. Carlos Cortés 0001, Irene Viola 0001, Jesús Gutiérrez 0001, Jack Jansen 0001, Shishir Subramanyam, Evangelos Alexiou, Pablo Pérez 0001, Narciso García, Pablo César |
ACM Trans. Multim. Comput. Commun. Appl. | 9 |
| 2024 | Designing and Evaluating a VR Lobby for a Socially Enriching Remote Opera Watching ExperienceabstractThe latest social VR technologies have enabled users to attend traditional media and arts performances together while being geographically removed, making such experiences accessible despite budget, distance, and other restrictions. In this work, we aim at improving the way remote performances are shared by designing and evaluating a VR theatre lobby which serves as a space for users to gather, interact, and relive the common experience of watching a virtual opera. We conducted an initial test with experts ($\mathrm{N}=10$, i.e., designers and opera enthusiasts) in pairs using our VR lobby prototype, developed based on the theoretical lobby design concept. A unique aspect of our experience is its highly realistic representation of users in the virtual space. The test results guided refinements to the VR lobby structure and implementation, aiming to improve the user experience and align it more closely with the social VR lobby's intended purpose. With the enhanced prototype, we ran a between-subject controlled study ($\mathrm{N}=40$) to compare the user experience in the social VR lobby between individuals and paired participants. To do so, we designed and validated a questionnaire to measure the user experience in the VR lobby. Results of our mixed-methods analysis, including interviews, questionnaire results, and user behavior, reveal the strength of our social VR lobby in connecting with other users, consuming the opera in a deeper manner, and exploring new possibilities beyond what is common in real life. All supplemental materials are available at https://github.com/cwi-dis/IEEEVR2024-VRLobby. Sueyoon Lee, Irene Viola 0001, Silvia Rossi 0001, Zhirui Guo, Ignacio Reimat, Kinga Lawicka, Alina Striner, Pablo César |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2023 | Affective Driver-Pedestrian Interaction: Exploring Driver Affective Responses toward Pedestrian Crossing Actions using Camera and Physiological SensorsabstractEliciting and capturing drivers’ affective responses in a realistic outdoor setting with pedestrians poses a challenge when designing in-vehicle, empathic interfaces. To address this, we designed a controlled, outdoor car driving circuit where drivers (N=27) drove and encountered pedestrian confederates who performed non-verbal positive or non-positive road crossing actions towards them. Our findings reveal that drivers reported higher valence upon observing positive, non-verbal crossing actions, and higher arousal upon observing non-positive crossing actions. Drivers’ heart signals (BVP, IBI and BPM), skin conductance and facial expressions (brow lowering, eyelid tightening, nose wrinkling, and lip stretching) all varied significantly when observing positive and non-positive actions. Our car driving study, by drawing on realistic driving conditions, further contributes to the development of in-vehicle empathic interfaces that leverage behavioural and physiological sensing. Through automatic inference of driver affect resulting from pedestrian actions, our work can enable novel empathic interfaces for supporting driver emotion self-regulation. Shruti Rao, Sabrina Wirjopawiro, Gerard Pons 0002, Thomas Röggla, Pablo César, Abdallah El Ali |
AutomotiveUI | 5 |
| 2023 | QAVA-DPC: Eye-Tracking Based Quality Assessment and Visual Attention Dataset for Dynamic Point Cloud in 6 DoFabstractPerceptual quality assessment of Dynamic Point Cloud (DPC) contents plays an important role in various Virtual Reality (VR) applications that involve human beings as the end user, understanding and modeling perceptual quality assessment is greatly enriched by insights from visual attention. However, incorporating aspects of visual attention in DPC quality models is largely unexplored, as ground-truth visual attention data is scarcely available. This paper presents a dataset containing subjective opinion scores and visual attention maps of DPCs, collected in a VR environment using eye-tracking technology. The data was collected during a subjective quality assessment experiment, in which subjects were instructed to watch and rate DPCs at various degradation levels under 6 degrees-of-freedom inspection, using a head-mounted display. The dataset comprises 5 reference DPC contents, with each reference encoded at 3 distortion levels using 3 different codecs, amounting to a total of 9 degraded DPC contents. Moreover, it includes 1,000 gaze trials from 40 participants, resulting in 15,000 visual attention maps in total. The curated dataset can serve as authentic benchmark data for assessing the performance of objective DPC quality metrics. Additionally, it establishes a link between quality assessment and visual attention within the context of DPC. This work deepens our understanding of DPC quality and visual attention, driving progress in the realm of VR experiences and perception. Xuemei Zhou, Irene Viola 0001, Evangelos Alexiou, Jack Jansen 0001, Pablo César |
ISMAR | 5 |
| 2023 | Extending 3-DoF Metrics to Model User Behaviour Similarity in 6-DoF Immersive ApplicationsabstractImmersive reality technologies, such as Virtual and Augmented Reality, have ushered a new era of user-centric systems, in which every aspect of the coding-delivery-rendering chain is tailored to the interaction of the users. Understanding the actual interactivity and behaviour of the users is still an open challenge and a key step to enabling such a user-centric system. Our main goal is to extend the applicability of existing behavioural methodologies for studying user navigation in the case of 6 Degree-of-Freedom (DoF). Specifically, we first compare the navigation in 6-DoF with its 3-DoF counterpart highlighting the main differences and novelties. Then, we define new metrics aimed at better modelling behavioural similarities between users in a 6-DoF system. We validate and test our solutions on real navigation paths of users interacting with dynamic volumetric media in 6-DoF Virtual Reality conditions. Our results show that metrics that consider both user position and viewing direction better perform in detecting user similarity while navigating in a 6-DoF system. Having easy-to-use but robust metrics that underpin multiple tools and answer the question "how do we detect if two users look at the same content?" open the gate to new solutions for a user-centric system. Silvia Rossi 0001, Irene Viola 0001, Laura Toni, Pablo César |
MMSys | 4 |
| 2023 | Evaluation of point cloud features for no-reference visual quality assessmentabstractThe development and widespread adoption of immersive XR applications has led to a renewed interest in representations that are capable of reproducing real-world objects and scenes with high fidelity. Among such representations, point clouds have attracted the interest of industry and academia alike, and new compression solutions have been developed to facilitate their adoption in mainstream applications. To ensure the best quality of experience for the end-user in limited bandwidth scenarios, new full-reference objective quality metrics have been proposed, promoting features designed specifically for point cloud contents. However, the performance of such features to predict the quality of point cloud contents when the reference is not available is largely unexplored. In this paper, we evaluate the performance of features commonly used to model point cloud distortions in a no-reference framework. The obtained features are integrated into a quality value through a support vector regression model. Results demonstrate the potential of full-reference features for no-reference assessment. Gwennan Smitskamp, Irene Viola 0001, Pablo César |
QoMEX | 3 |
| 2023 | From Video to Hybrid Simulator: Exploring Affective Responses toward Non-Verbal Pedestrian Crossing Actions Using Camera and Physiological SensorsabstractCapturing drivers’ affective responses given driving context and driver-pedestrian interactions remains a challenge for designing in-vehicle, empathic interfaces. To address this, we conducted two lab-based studies using camera and physiological sensors. Our first study collected participants’ (N = 21) emotion self-reports and physiological signals (including facial temperatures) toward non-verbal, pedestrian crossing videos from the Joint Attention for Autonomous Driving dataset. Our second study increased realism by employing a hybrid driving simulator setup to capture participants’ affective responses (N = 24) toward enacted, non-verbal pedestrian crossing actions. Key findings showed: (a) non-positive actions in videos elicited higher arousal ratings, whereas different in-video pedestrian crossing actions significantly influenced participants’ physiological signals. (b) Non-verbal pedestrian interactions in the hybrid simulator setup significantly influenced participants’ facial expressions, but not their physiological signals. We contribute to the development of in-vehicle empathic interfaces that draw on behavioral and physiological sensing to in-situ infer driver affective responses during non-verbal pedestrian interactions. Shruti Rao, Surjya Ghosh, Gerard Pons 0002, Thomas Röggla, Pablo César, Abdallah El Ali |
Int. J. Hum. Comput. Interact. | 5 |
| 2023 | Group Synchrony for Emotion Recognition Using Physiological SignalsabstractDuring group interactions, we react and modulate our emotions and behaviour to the group through phenomena including emotion contagion and physiological synchrony. Previous work on emotion recognition through video/image has shown that group context information improves the classification performance. However, when using physiological data, literature mostly focuses on intrapersonal models that leave-out group information, while interpersonal models are unexplored. This paper introduces a new interpersonal Weighted Group Synchrony approach, which relies on Electrodermal Activity (EDA) and Heart-Rate Variability (HRV). We perform an analysis of synchrony metrics applied across diverse data representations (EDA and HRV morphology and features, recurrence plot, spectrogram), to identify which metrics and modalities better characterise physiological synchrony for emotion recognition. We explored two datasets (AMIGOS and K-EmoCon), covering different group sizes (4 vs dyad) and group-based activities (video-watching vs conversation). The experimental results show that integrating group information improves arousal and valence classification, across all datasets, with the exception of K-EmoCon on valence. The proposed method was able to attain mean M-F1 of$\approx$72.15% arousal and 81.16% valence for AMIGOS, and M-F1 of$\approx$52.63% arousal, 65.09% valence for K-EmoCon, surpassing previous work results for K-EmoCon on arousal, and providing a new baseline on AMIGOS for long-videos. Patrícia J. Bota, Tianyi Zhang 0013, Abdallah El Ali, Ana Fred, Hugo Silva 0001, Pablo César |
IEEE Trans. Affect. Comput. | 6 |
| 2023 | Weakly-Supervised Learning for Fine-Grained Emotion Recognition Using Physiological SignalsabstractInstead of predicting just one emotion for one activity (e.g., video watching), fine-grained emotion recognition enables more temporally precise recognition. Previous works on fine-grained emotion recognition require segment-by-segment, fine-grained emotion labels to train the recognition algorithm. However, experiments to collect these labels are costly and time-consuming compared with only collecting one emotion label after the user watched that stimulus (i.e., the post-stimuli emotion labels). To recognize emotions at a finer granularity level when trained with only post-stimuli labels, we propose an emotion recognition algorithm based on Deep Multiple Instance Learning (EDMIL) using physiological signals.EDMILrecognizes fine-grained valence and arousal (V-A) labels by identifying which instances represent the post-stimuli V-A annotated by users after watching the videos. Instead of fully-supervised training, the instances are weakly-supervised by the post-stimuli labels in the training stage. The V-A of instances are estimated by the instance gains, which indicate the probability of instances to predict the post-stimuli labels. We testedEDMILon three different datasets,CASE,MERCAandCEAP-360VR, collected in three different environments: desktop, mobile and HMD-based Virtual Reality, respectively. Recognition results validated with the fine-grained V-A self-reports show that for subject-independent 3-class classification (high/neutral/low),EDMILobtains promising recognition accuracies: 75.63% and 79.73% for V-A onCASE, 70.51% and 67.62% for V-A onMERCAand 65.04% and 67.05% for V-A onCEAP-360VR. Our ablation study shows that all components ofEDMILcontribute to both the classification and regression tasks. Our experiments also show that (1) compared with fully-supervised learning, weakly-supervised learning can reduce the problem of overfitting caused by the temporal mismatch between fine-grained annotations and physiological signals, (2) instance segment lengths between 1-2 s result in the highest recognition accuracies and (3)EDMILperforms best if post-stimuli annotations consist of less than 30% or more than 60% of the entire video watching. Tianyi Zhang 0013, Abdallah El Ali, Chen Wang 0034, Alan Hanjalic, Pablo César |
IEEE Trans. Affect. Comput. | 5 |
| 2023 | Multimodal-Based and Aesthetic-Guided Narrative Video SummarizationabstractNarrative videos usually illustrate the main content through multiple narrative information such as audios, video frames and subtitles. Existing video summarization approaches rarely consider the multiple dimensional narrative inputs, or ignore the impact of shots artistic assembly when directly applied to narrative videos. This paper introduces a multimodal-based and aesthetic-guided narrative video summarization method. Our method leverages multimodal information including visual content, subtitles and audio information through our specified key shots selection, subtitle summarization, and highlight extraction components. Furthermore, under the guidance of cinematographic aesthetic, we design a novel shots assembly module to ensure the shot content completeness and then assemble the selected shots into a desired summary. Besides, our method also provides the flexible specification for shots selection, to achieve which it automatically selects semantically related shots according to the user-designed text. By conducting a large number of quantitative experimental evaluations and user studies, we demonstrate that our method effectively preserves important narrative information of the original video, and it is capable of rapidly producing high-quality and aesthetic-guided narrative video summaries. Jiehang Xie, Xuanbai Chen, Tianyi Zhang 0013, Shao-Ping Lu, Pablo César, Yulu Yang |
IEEE Trans. Multim. | 6 |
| 2023 | CEAP-360VR: A Continuous Physiological and Behavioral Emotion Annotation Dataset for 360$^\circ$ VR VideosabstractWatching 360$^\circ$videos using Virtual Reality (VR) head-mounted displays (HMDs) provides interactive and immersive experiences, where videos can evoke different emotions. Existing emotion self-report techniques within VR however are either retrospective or interrupt the immersive experience. To address this, we introduce theContinuous Physiological and Behavioral Emotion Annotation Dataset for 360$^\circ$Videos (CEAP-360VR). We conducted a controlled study (N=32) where participants used a Vive Pro Eye HMD to watch eight validated affective 360$^\circ$video clips, and annotated their valence and arousal (V-A) continuously. We collected (a) behavioral (head and eye movements; pupillometry) signals (b) physiological (heart rate, skin temperature, electrodermal activity) responses (c) momentary emotion self-reports (d) within-VR discrete emotion ratings (e) motion sickness, presence, and workload. We show the consistency of continuous annotation trajectories and verify their mean V-A annotations. We find high consistency between viewed 360$^\circ$video regions across subjects, with higher consistency for eye than head movements. We furthermore run baseline classification experiments, where Random Forest classifiers with 2s segments show good accuracies for subject-independent models: 66.80% (V) and 64.26% (A) for binary classification; 49.92% (V) and 52.20% (A) for 3-class classification. Our open dataset allows further experiments with continuous emotion self-reports collected in 360$^\circ$VR environments, which can enable automatic assessment of immersive Quality of Experience (QoE) andmomentary affective states. Tong Xue, Abdallah El Ali, Tianyi Zhang 0013, Pablo César |
IEEE Trans. Multim. | 5 |
| 2023 | Few-Shot Learning for Fine-Grained Emotion Recognition Using Physiological SignalsabstractFine-grained emotion recognition can model the temporal dynamics of emotions, which is more precise than predicting one emotion retrospectively for an activity (e.g., video clip watching). Previous works require large amounts of continuously annotated data to train an accurate recognition model, however experiments to collect such large amounts of continuously annotated physiological signals are costly and time-consuming. To overcome this challenge, we propose an Emotion recognition algorithm based on Deep Siamese Networks (EmoDSN) which can rapidly converge on a small amount of training data, typically less than 10 samples per class (i.e., <10 shot). EmoDSN recognizes fine-grained valence and arousal (V-A) labels by maximizing the distance metric between signal segments with different V-A labels. We tested EmoDSN on three different datasets collected in three different environments: desktop, mobile and HMD-based virtual reality, respectively. The results from our experiments show that EmoDSN achieves promising results for both one-dimension binary (high/low V-A, 1D-2 C) and two-dimensional 5-class (four quadrants of V- A space + neutral, 2D-5 C) classification. We get an averaged accuracy of 76.04, 76.62 and 57.62% for 1D-2 C valence, 1D-2 C arousal, and 2D-5 C, respectively, by using only 5 shots of training data. Our experiments show that EmoDSN can achieve better results if we select training samples from the changing points of emotion or the ending moments of video watching. Tianyi Zhang 0013, Abdallah El Ali, Alan Hanjalic, Pablo César |
IEEE Trans. Multim. | 4 |
| 2023 | Is that My Heartbeat? Measuring and Understanding Modality-Dependent Cardiac Interoception in Virtual RealityabstractMeasuring interoception ('perceiving internal bodily states') has diagnostic and wellbeing implications. Since heartbeats are distinct and frequent, various methods aim at measuring cardiac interoceptive accuracy (CIAcc). However, the role of exteroceptive modalities for representing heart rate (HR) across screen-based and Virtual Reality (VR) environments remains unclear. Using a PolarH10 HR monitor, we develop a modality-dependent cardiac recognition task that modifies displayed HR. In a mixed-factorial design (N=50), we investigate how task environment (Screen, VR), modality (Audio, Visual, Audio-Visual), and real-time HR modifications (±15%, ±30%, None) influence CIAcc, interoceptive awareness, mind-body measures, VR presence, and post-experience responses. Findings showed that participants confused their HR with underestimates up to 30%; environment did not affect CIAcc but influenced mind-related measures; modality did not influence CIAcc, however including audio increased interoceptive awareness; and VR presence inversely correlated with CIAcc. We contribute a lightweight and extensible cardiac interoception measurement method, and implications for biofeedback displays. Abdallah El Ali, Rayna Ney, Zeph M. C. van Berlo, Pablo César |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Understanding and Designing Avatar Biosignal Visualizations for Social Virtual Reality EntertainmentabstractVisualizing biosignals can be important for social Virtual Reality (VR), where avatar non-verbal cues are missing. While several biosignal representations exist, designing effective visualizations and understanding user perceptions within social VR entertainment remains unclear. We adopt a mixed-methods approach to design biosignals for social VR entertainment. Using survey (N=54), context-mapping (N=6), and co-design (N=6) methods, we derive four visualizations. We then ran a within-subjects study (N=32) in a virtual jazz-bar to investigate how heart rate (HR) and breathing rate (BR) visualizations, and signal rate, influence perceived avatar arousal, user distraction, and preferences. Findings show that skeuomorphic visualizations for both biosignals allow differentiable arousal inference; skeuomorphic and particles were least distracting for HR, whereas all were similarly distracting for BR; biosignal perceptions often depend on avatar relations, entertainment type, and emotion inference of avatars versus spaces. We contribute HR and BR visualizations, and considerations for designing social VR entertainment biosignal visualizations. Sueyoon Lee, Abdallah El Ali, Maarten Wijntjes, Pablo César |
CHI | 4 |
| 2022 | Digital Proxemics: Designing Social and Collaborative Interaction in Virtual EnvironmentsabstractBehaviour in virtual environments might be informed by our experiences in physical environments, but virtual environments are not constrained by the same physical, perceptual, or social cues. Instead of replicating the properties of physical spaces, one can create virtual experiences that diverge from reality by dynamically manipulating environmental, aural, and social properties. This paper explores digital proxemics, which describe how we use space in virtual environments and how the presence of others influences our behaviours, interactions, and movements. First, we frame the open challenges of digital proxemics in terms of activity, social signals, audio design, and environment. We explore a subset of these challenges through an evaluation that compares two audio designs and two displays with different social signal affordances: head-mounted display (HMD) versus desktop PC. We use quantitative methods using instrumented tracking to analyse behaviour, demonstrating how personal space, proximity, and attention compare between desktop PC and HMDs. Julie R. Williamson, Joseph O'Hagan, John Alexis Guerra Gómez, John Williamson 0001, Pablo César, David A. Shamma |
CHI | 5 |
| 2022 | Impact of Self-View Latency on Quality of Experience: Analysis of Natural Interaction in XR EnvironmentsabstractThe rise of eXtended Reality (XR) has led to multiple ways of including the user’s body in interactive experiences. However, the delay limits of self-view rendering in interactive XR remain unexplored. This article presents a minimum self-view latency system and an interactive task-based experiment to study the influence of different levels of self-view delay on Quality of Experience (QoE) and task performance. During the experiment, 23 users tested 8 delay conditions (from 190 to 597 ms) while building block-based models. The results show a hard threshold in terms of involvement and overall quality around 450ms. However, the impact on adaptation and execution time was less pronounced. This indicates that although users adapted to the task in a certain way, their immersion was severely affected above a certain self-view delay value. Carlos Cortés 0001, Jesús Gutiérrez 0001, Pablo Pérez 0001, Irene Viola 0001, Pablo César, Narciso García |
ICIP | 5 |
| 2022 | Mediascape XR: A Cultural Heritage Experience in Social VRabstractSocial virtual reality (VR) allows multiple remote users to interact in a shared space, unveiling new possibilities for communication in immersive environments. Mediascape XR presents a social VR experience that teleports 3D representations of remote users, using volumetric video, to a virtual museum. It enables visitors to interact with cultural heritage artifacts while allowing social interactions in real time between them. The application is designed following a human-centered approach, enabling an interactive, educating, and entertaining experience. Ignacio Reimat, Yanni Mei, Evangelos Alexiou, Jack Jansen 0001, Jie Li 0064, Shishir Subramanyam, Irene Viola 0001, Johan Oomen, Pablo César |
ACM Multimedia | 9 |
| 2022 | Evaluating the Impact of Tiled User-Adaptive Real-Time Point Cloud Streaming on VR Remote CommunicationabstractRemote communication has rapidly become a part of everyday life in both professional and personal contexts. However, popular video conferencing applications present limitations in terms of quality of communication, immersion and social meaning. VR remote communication applications offer a greater sense of co-presence and mutual sensing of emotions between remote users. Previous research on these applications has shown that realistic point cloud user reconstructions offer better immersion and communication as compared to synthetic user avatars. However, photorealistic point clouds require a large volume of data per frame and are challenging to transmit over bandwidth-limited networks. Recent research has demonstrated significant improvements to perceived quality by optimizing the usage of bandwidth based on the position and orientation of the user's viewport with user-adaptive streaming. In this work, we developed a real-time VR communication application with an adaptation engine that features tiled user-adaptive streaming based on user behaviour. The application also supports traditional network adaptive streaming. The contribution of this work is to evaluate the impact of tiled user-adaptive streaming on quality of communication, visual quality, system performance and task completion in a functional live VR remote communication system. We performed a subjective evaluation with 33 users to compare the different streaming conditions with a neck exercise training task. As a baseline, we use uncompressed streaming requiring approximately 300 megabits per second and our solution achieves similar visual quality with tiled adaptive streaming at 14 megabits per second. We also demonstrate statistically significant gains in the quality of interaction and improvements to system performance and CPU consumption with tiled adaptive streaming as compared to the more traditional network adaptive streaming. Shishir Subramanyam, Irene Viola 0001, Jack Jansen 0001, Evangelos Alexiou, Alan Hanjalic, Pablo César |
ACM Multimedia | 6 |
| 2022 | Subjective QoE Evaluation of User-Centered Adaptive Streaming of Dynamic Point CloudsabstractTechnological advances in head-mounted displays and novel real-time 3D acquisition and reconstruction solutions have fostered the development of 6 Degrees of Freedom (6DoF) teleimmersive systems for social VR applications. Point clouds have emerged as a popular format for such applications, owing to their simplicity and versatility; yet, dense point cloud contents are too large to deliver directly over bandwidth-limited networks. In this context, user-adaptive delivery mechanisms are a promising solution to exploit the increased range of motion offered by 6DoF VR applications to yield gains in perceived quality of 3D point cloud user representations, while reducing their bandwidth requirements. In this paper, we perform a user study in VR to quantify the gains adaptive tile selection strategies can bring with respect to non-adaptive solutions. In particular, we define an auxiliary utility function, we employ established methods from the literature and newly-proposed schemes for distributing the bit budget across the tiles, and we evaluate them together with non-adaptive streaming baselines through subjective QoE assessment. Results confirm that considerable gains can be obtained with user-adaptive streaming, achieving bit rate gains of up to 65% with respect to a non-adaptive approach to deliver comparable quality. Our analysis provides useful insights for the design and development of social VR applications. Shishir Subramanyam, Irene Viola 0001, Jack Jansen 0001, Evangelos Alexiou, Alan Hanjalic, Pablo César |
QoMEX | 6 |
| 2022 | Designing Real-time, Continuous QoE Score Acquisition Techniques for HMD-based 360°VR Video WatchingabstractWatching HMD-based 360° video has become in-creasing popular as a medium for immersive viewing of photo-realistic content. To evaluate subjective video quality, researchers typically prompt users to provide an overall Quality of Experience (QoE) score after viewing a stimulus. However, since users can adjust their viewport throughout a 360° video, a higher level of spatiotemporal granularity is needed for adaptive 360° video streaming. To address this, we design several real-time, continuous QoE annotation input and peripheral visualization techniques, with the goal of minimizing mental workload and distraction during score acquisition. Drawing on two parallel co-design sessions with seven experts, we find that touchpad and joystick are most suitable for continuous input, with DotMorph (circle with tick label that varies in filling) for peripheral state feedback. We contribute design findings for testing QoE score acquisition techniques during HMD-based 360° video watching, which enable more precise optimization of adaptive video streaming quality. Tong Xue, Abdallah El Ali, Irene Viola 0001, Pablo César |
QoMEX | 5 |
| 2022 | Designing a VR Lobby for Remote Opera Social ExperiencesabstractSeveral social VR platforms support virtual entertainment events, however their value for post-show activities remains unclear. Through a user-centered approach, we design a social VR lobby experience to enrich four motivations of theatre-goers: social, intellectual, emotional, and spiritual engagement. We ran a context-mapping focus group session with professionals (N=6) to conceptualize the social VR space for digital opera experiences. Based on our findings, we propose a social VR lobby consisting of four rooms: 1) a Bar for social engagement, 2) an Info Booth for intellectual engagement, 3) a Photo Zone for emotional engagement, and 4) an Interactive Stage for spiritual engagement. Based on this work, we plan to experimentally evaluate audience experiences in each room in order to create a social VR lobby template for theater experiences. Sueyoon Lee, Alina Striner, Pablo César |
IMX | 3 |
| 2022 | The Co-Creation Space: An Online Safe Space for Community Opera CreationabstractThis work presents the Co-Creation Space, a multilingual platform for professional and community artists to 1) generate raw artistic ideas, and 2) discuss and reflect on the shared meaning of those ideas. The paper describes the architecture and the technology behind the platform, and how it was used to facilitate the communication process during several user trials. By supporting ideation sessions around media items guided by a facilitator and allowing users to express themselves and be part of the creation of an artistic product, participants were enabled to access new cultural spaces and be part of the creative process. Thomas Röggla, Alina Striner, Héctor Rivas Pagador, Pablo César |
IMX | 4 |
| 2022 | Subjective Evaluation of Visual Quality and Simulator Sickness of Short 360$^\circ$ Videos: ITU-T Rec. P.919abstractRecently an impressive development in immersive technologies, such as Augmented Reality (AR), Virtual Reality (VR) and 360${^\circ }$video, has been witnessed. However, methods for quality assessment have not been keeping up. This paper studies quality assessment of 360${^\circ }$video from the cross-lab tests (involving ten laboratories and more than 300 participants) carried out by the Immersive Media Group (IMG) of the Video Quality Experts Group (VQEG). These tests were addressed to assess and validate subjective evaluation methodologies for 360${^\circ }$video. Audiovisual quality, simulator sickness symptoms, and exploration behavior were evaluated with short (from 10 seconds to 30 seconds) 360${^\circ }$sequences. The following factors’ influences were also analyzed: assessment methodology, sequence duration, Head-Mounted Display (HMD) device, uniform and non-uniform coding degradations, and simulator sickness assessment methods. The obtained results have demonstrated the validity of Absolute Category Rating (ACR) and Degradation Category Rating (DCR) for subjective tests with 360${^\circ }$videos, the possibility of using 10-second videos (with or without audio) when addressing quality evaluation of coding artifacts, as well as any commercial HMD (satisfying minimum requirements). Also, more efficient methods than the long Simulator Sickness Questionnaire (SSQ) have been proposed to evaluate related symptoms with 360${^\circ }$videos. These results have been instrumental for the development of the ITU-T Recommendation P.919. Finally, the annotated dataset from the tests is made publicly available for the research community. Jesús Gutiérrez 0001, Pablo Pérez 0001, Marta Orduna, Ashutosh Singla, Carlos Cortés 0001, Pramit Mazumdar, Irene Viola 0001, Kjell Brunnström, Federica Battisti, Natalia Cieplinska, Dawid Juszka, Lucjan Janowski, Mikolaj Leszczuk, Anthony Adeyemi-Ejeye, Yaosi Hu, Zhenzhong Chen 0001, Glenn Van Wallendael, Peter Lambert, César Díaz, John Hedlund, Omar Hamsis, Stephan Fremerey, Frank Hofmeyer, Alexander Raake, Pablo César, Marco Carli, Narciso García |
IEEE Trans. Multim. | 25 |
| 2021 | CakeVR: A Social Virtual Reality (VR) Tool for Co-designing CakesabstractCake customization services allow clients to collaboratively personalize cakes with pastry chefs. However, remote (e.g., email) and in-person co-design sessions are prone to miscommunication, due to natural restrictions in visualizing cake size, decoration, and celebration context. This paper presents the design, implementation, and expert evaluation of a social VR application (CakeVR) that allows a client to remotely co-design cakes with a pastry chef, through real-time realistic 3D visualizations. Drawing on expert semi-structured interviews (4 clients, 5 pastry chefs), we distill and incorporate 8 design requirements into our CakeVR prototype. We evaluate CakeVR with 10 experts (6 clients, 4 pastry chefs) using cognitive walkthroughs, and find that it supports ideation and decision making through intuitive size manipulation, color/flavor selection, decoration design, and custom celebration theme fitting. Our findings provide recommendations for enabling co-design in social VR and highlight CakeVR’s potential to transform product design communication through remote interactive and immersive co-design. Yanni Mei, Jie Li 0064, Huib de Ridder, Pablo César |
CHI | 4 |
| 2021 | Proxemics and Social Interactions in an Instrumented Virtual Reality WorkshopabstractVirtual environments (VEs) can create collaborative and social spaces, which are increasingly important in the face of remote work and travel reduction. Recent advances, such as more open and widely available platforms, create new possibilities to observe and analyse interaction in VEs. Using a custom instrumented build of Mozilla Hubs to measure position and orientation, we conducted an academic workshop to facilitate a range of typical workshop activities. We analysed social interactions during a keynote, small group breakouts, and informal networking/hallway conversations. Our mixed-methods approach combined environment logging, observations, and semi-structured interviews. The results demonstrate how small and large spaces influenced group formation, shared attention, and personal space, where smaller rooms facilitated more cohesive groups while larger rooms made small group formation challenging but personal space more flexible. Beyond our findings, we show how the combination of data and insights can fuel collaborative spaces’ design and deliver more effective virtual workshops. Julie R. Williamson, Jie Li 0064, Vinoba Vinayagamoorthy, David A. Shamma, Pablo César |
CHI | 5 |
| 2021 | RCEA-360VR: Real-time, Continuous Emotion Annotation in 360° VR Videos for Collecting Precise Viewport-dependent Ground Truth LabelsabstractPrecise emotion ground truth labels for 360° virtual reality (VR) video watching are essential for fine-grained predictions under varying viewing behavior. However, current annotation techniques either rely on post-stimulus discrete self-reports, or real-time, continuous emotion annotations (RCEA) but only for desktop/mobile settings. We present RCEA for 360° VR videos (RCEA-360VR), where we evaluate in a controlled study (N=32) the usability of two peripheral visualization techniques: HaloLight and DotSize. We furthermore develop a method that considers head movements when fusing labels. Using physiological, behavioral, and subjective measures, we show that (1) both techniques do not increase users’ workload, sickness, nor break presence (2) our continuous valence and arousal annotations are consistent with discrete within-VR and original stimuli ratings (3) users exhibit high similarity in viewing behavior, where fused ratings perfectly align with intended labels. Our work contributes usable and effective techniques for collecting fine-grained viewport-dependent emotion labels in 360° VR. Tong Xue, Abdallah El Ali, Tianyi Zhang 0013, Pablo César |
CHI | 5 |
| 2021 | A New Challenge: Behavioural Analysis Of 6-DOF User When Consuming Immersive MediaabstractThanks to recent advances in computer graphics, wearable technology, and connectivity, Virtual Reality (VR) has landed in our daily life. A key novelty in VR is the role of the user, which has turned from merely passive to entirely active. Thus, improving any aspect of the coding-delivery-rendering chain starts with the need for understanding user behaviour. To do so, we investigate the navigation trajectories of users within a 6-Degrees-of-Freedom (DoF) VR environment. Specifically, we investigate the main differences and similarities between 3 and 6-DoF navigation through existing methodologies adopted to study user behaviour in 3-DoF settings. Our simulation results, based on real navigation paths of users while displaying dynamic volumetric media in 6-DoF conditions, show the limitations of clustering algorithms for 3-DoF in assessing user similarity in 6-DoF. Given these observations, we state the need for developing new solutions for the analysis of 6-DoF trajectories. Silvia Rossi 0001, Irene Viola 0001, Laura Toni, Pablo César |
ICIP | 4 |
| 2021 | Evaluating the user Experience of a Photorealistic Social VR MovieabstractWe all enjoy watching movies together. However, this is not always possible if we live apart. While we can remotely share our screens, the experience differs from being together. We present a social Virtual Reality (VR) system that captures, reconstructs, and transmits multiple users’ volumetric representations into a commercially produced 3D virtual movie, so they have the feeling of “being there” together. We conducted a 48-user experiment where we invited users to experience the virtual movie either using a Head Mounted Display (HMD) or using a 2D screen with a game controller. In addition, we invited 14 VR experts to experience both the HMD and the screen version of the movie and discussed their experiences in two focus groups. Our results showed that both end-users and VR experts found that the way they navigated and interacted inside a 3D virtual movie was novel. They also found that the photorealistic volumetric representations enhanced feelings of co-presence. Our study lays the groundwork for future interactive and immersive VR movie co-watching experiences. Jie Li 0064, Shishir Subramanyam, Jack Jansen 0001, Yanni Mei, Ignacio Reimat, Kinga Lawicka, Pablo César |
ISMAR | 7 |
| 2021 | CWIPC-SXR: Point Cloud dynamic human dataset for Social XRabstractReal-time, immersive telecommunication systems are quickly becoming a reality, thanks to the advances in acquisition, transmission, and rendering technologies. Point clouds in particular serve as a promising representation in these type of systems, offering photorealistic rendering capabilities with low complexity. Further development of transmission, coding, and quality evaluation algorithms, though, is currently hindered by the lack of publicly available datasets that represent realistic scenarios of remote communication between people in real-time. In this paper, we release a dynamic point cloud dataset that depicts humans interacting in social XR settings. Using commodity hardware, we capture a total of 45 unique sequences, according to several use cases for social XR. As part of our release, we provide annotated raw material, resulting point cloud sequences, and an auxiliary software toolbox to acquire, process, encode, and visualize data, suitable for real-time applications. The dataset can be accessed via the following link: https://www.dis.cwi.nl/cwipc-sxr-dataset/. Ignacio Reimat, Evangelos Alexiou, Jack Jansen 0001, Irene Viola 0001, Shishir Subramanyam, Pablo César |
MMSys | 6 |
| 2021 | Co-creation Stage: a Web-based Tool for Collaborative and Participatory Co-located Art PerformancesabstractIn recent years, artists and communities have expressed the desire to work with tools that facilitate co-creation and allow distributed community performances. These performances can be spread over several physical stages, connecting them on real-time towards a single experience with the audience distributed along them. This enables a wider remote audience consuming the performance through their own devices, and even grants the participation of remote users in the show. In this paper we introduce the Co-creation Stage, a web-based tool that allows managing heterogeneous content sources, with a particular focus on live and on-demand media, across several distributed devices. The Co-creation Stage is part of the toolset developed in the Traction H2020 project which enables community performing art shows, where professional artists and non-professional participants perform together from different stages and locations. Here we present the design process, the architecture and the main functionalities of the tool as well as the results of the first user evaluation with Opera houses and artists. Héctor Rivas Pagador, Ana Dominguez 0001, Stefano Masneri, Iñigo Tamayo, Mikel Zorrilla, Pedro Almeida 0002, Jie Li 0064, Alina Striner, Pablo César |
IMX | 9 |
| 2021 | Towards Immersive and Social Audience Experience in Remote VR OperaabstractOpera is a historic art that struggles to be approachable to modern audiences. In partnership with the Irish National Opera (INO), this work considers how VR may be used to develop a new form of immersive opera. To this end, we ran three open-ended focus groups to consider how creative, multisensory, and social VR technology may be employed in digital opera. Our findings assert the importance of creating an immersive experience by safely giving audiences agency to interact, to democratize personal and social experiences, and to consider different ways of representing their bodies, their social rituals, and the virtual social space. Using these findings, we envision a new form of VR opera that couples physical traditions with digital affordances. Alina Striner, Sarah Halpin, Thomas Röggla, Pablo César |
IMX | 4 |
| 2020 | ThermalWear: Exploring Wearable On-chest Thermal Displays to Augment Voice Messages with AffectabstractVoice is a rich modality for conveying emotions, however emotional prosody production can be situationally or medically impaired. Since thermal displays have been shown to evoke emotions, we explore how thermal stimulation can augment perception of neutrally-spoken voice messages with affect. We designed ThermalWear, a wearable on-chest thermal display, then tested in a controlled study (N=12) the effects of fabric, thermal intensity, and direction of change. Thereafter, we synthesized 12 neutrally-spoken voice messages, validated (N=7) them, then tested (N=12) if thermal stimuli can augment their perception with affect. We found warm and cool stimuli (a) can be perceived on the chest, and quickly without fabric (4.7-5s) (b) do not incur discomfort (c) generally increase arousal of voice messages and (d) increase / decrease message valence, respectively. We discuss how thermal displays can augment voice perception, which can enhance voice assistants and support individuals with emotional prosody impairments. Abdallah El Ali, Swamy Ananthanarayan, Thomas Röggla, Jack Jansen 0001, Jessica Hartcher-O'Brien, Kaspar M. B. Jansen, Pablo César |
CHI | 8 |
| 2020 | RCEA: Real-time, Continuous Emotion Annotation for Collecting Precise Mobile Video Ground Truth LabelsabstractCollecting accurate and precise emotion ground truth labels for mobile video watching is essential for ensuring meaningful predictions. However, video-based emotion annotation techniques either rely on post-stimulus discrete self-reports, or allow real-time, continuous emotion annotations (RCEA) only for desktop settings. Following a user-centric approach, we designed an RCEA technique for mobile video watching, and validated its usability and reliability in a controlled, indoor (N=12) and later outdoor (N=20) study. Drawing on physiological measures, interaction logs, and subjective workload reports, we show that (1) RCEA is perceived to be usable for annotating emotions while mobile video watching, without increasing users' mental workload (2) the resulting time-variant annotations are comparable with intended emotion attributes of the video stimuli (classification error for valence: 8.3%; arousal: 25%). We contribute a validated annotation technique and associated annotation fusion method, that is suitable for collecting fine-grained emotion annotations while users watch mobile videos. Tianyi Zhang 0013, Abdallah El Ali, Chen Wang 0034, Alan Hanjalic, Pablo César |
CHI | 5 |
| 2020 | User Centered Adaptive Streaming of Dynamic Point Clouds with Low Complexity TilingabstractIn recent years, the development of devices for acquisition and rendering of 3D contents have facilitated the diffusion of immersive virtual reality experiences. In particular, the point cloud representation has emerged as a popular format for volumetric photorealistic reconstructions of dynamic real world objects, due to its simplicity and versatility. To optimize the delivery of the large amount of data needed to provide these experiences, adaptive streaming over HTTP is a promising solution. In order to ensure the best quality of experience within the bandwidth constraints, adaptive streaming is combined with tiling to optimize the quality of what is being visualized by the user at a given moment; as such, it has been successfully used in the past for omnidirectional contents. However, its adoption to the point cloud streaming scenario has only been studied to optimize multi-object delivery. In this paper, we present a low-complexity tiling approach to perform adaptive streaming of point cloud content. Tiles are defined by segmenting each point cloud object in several parts, which are then independently encoded. In order to evaluate the approach, we first collect real navigation paths, obtained through a user study in 6 degrees of freedom with 26 participants. The variation in movements and interaction behaviour among users indicate that a user-centered adaptive delivery could lead to sensible gains in terms of perceived quality. Evaluation of the performance of the proposed tiling approach against state of the art solutions for point cloud compression, performed on the collected navigation paths, confirms that considerable gains can be obtained by exploiting user-adaptive streaming, achieving bitrate gains up to 57% with respect to a non-adaptive approach with the same codec. Moreover, we demonstrate that the selection of navigation data has an impact on the relative objective scores. Shishir Subramanyam, Irene Viola 0001, Alan Hanjalic, Pablo César |
ACM Multimedia | 4 |
| 2020 | A pipeline for multiparty volumetric video conferencing: transmission of point clouds over low latency DASHabstractThe advent of affordable 3D capture and display hardware is making volumetric videoconferencing feasible. This technology increases the immersion of the participants, breaking the flat restriction of 2D screens, by allowing them to collaborate and interact in shared virtual reality spaces. In this paper we introduce the design and development of an architecture intended for volumetric videoconferencing that provides a highly realistic 3D representation of the participants, based on pointclouds. A pointcloud representation is suitable for real-time applications like video conferencing, due to its low-complexity and because it does not need a time consuming reconstruction process. As transport protocol we selected low latency DASH, due to its popularity and client-based adaptation mechanisms for tiling. This paper presents the architectural design, details the implementation, and provides some referential results. The demo will showcase the system in action, enabling volumetric videoconferencing using pointclouds. Jack Jansen 0001, Shishir Subramanyam, Romain Bouqueau, Gianluca Cernigliaro, Marc Martos Cabré, Pablo César |
MMSys | 7 |
| 2020 | A Color-Based Objective Quality Metric for Point Cloud ContentsabstractIn recent years, point clouds have gained popularity as a promising representation for volumetric contents in immersive scenarios. Standardization bodies such as MPEG have been developing new compression standards for point cloud contents to reduce the volume of data, while maintaining an acceptable level of visual quality. To do so, reliable metrics are needed in order to automatically estimate the perceptual quality of degraded point cloud contents. Whereas several objective metrics have been developed to assess the geometrical impairment of degraded point cloud contents, fewer publications have been devoted to evaluating color artifacts. In this paper, we propose new color-based objective metrics for quality evaluation of point cloud contents. Our work extracts color statistics from both reference and degraded point cloud contents, in order to assess the level of impairment. Using publicly available ground-truth data, we compare the performance of our proposed work with state-of-the-art metrics, and we demonstrate how the color metrics are able to achieve comparable results with respect to widely adopted solutions. Moreover, we combine color- and geometry-based metrics in order to provide a global quality score. The novelty of our works resides in simultaneously taking both degradation types into account, while being independent of the rendering process. Results show that our solution is able to overcome the limitations of focusing on only one type of degradation, achieving better performance with respect to current metrics. Irene Viola 0001, Shishir Subramanyam, Pablo César |
QoMEX | 3 |
| 2020 | Comparing the Quality of Highly Realistic Digital Humans in 3DoF and 6DoF: A Volumetric Video Case StudyabstractVirtual Reality (VR) and Augmented Reality (AR) applications have seen a drastic increase in commercial popularity. Different representations have been used to create 3D reconstructions for AR and VR. Point clouds are one such representation characterized by their simplicity and versatility, making them suitable for real time applications, such as reconstructing humans for social virtual reality. In this study, we evaluate how the visual quality of digital humans, represented using point clouds, is affected by compression distortions. We compare the performance of the upcoming point cloud compression standard against an octree-based anchor codec. Two different VR viewing conditions enabling 3- and 6 degrees of freedom are tested, to understand how interacting in the virtual space affects the perception of quality. To the best of our knowledge, this is the first work performing user quality evaluation of dynamic point clouds in VR; in addition, contributions of the paper include quantitative data and empirical findings. Results highlight how perceived visual quality is affected by the tested content, and how current data sets might not be sufficient to comprehensively evaluate compression solutions. Moreover, shortcomings in how point cloud encoding solutions handle visually-lossless compression are discussed. Shishir Subramanyam, Jie Li 0064, Irene Viola 0001, Pablo César |
VR | 4 |
| 2020 | A Reduced Reference Metric for Visual Quality Evaluation of Point Cloud ContentsabstractPoint cloud representation has seen a surge of popularity in recent years, thanks to its capability to reproduce volumetric scenes in immersive scenarios. New compression solutions for streaming of point cloud contents have been proposed, which require objective quality metrics to reliably assess the level of degradation introduced by coding and transmission distortions. In this context, reduced reference metrics aim to predict the visual quality of the transmitted contents, while requiring only a small set of features to be sent in addition to the streamed media. In this paper, we propose a reduced reference metric to predict the quality of point cloud contents under compression distortions. To do so, we extract a small set of statistical features from the reference point cloud in the geometry, color and normal vector domain, which can be used at the receiver side to assess the visual degradation of the content. Using publicly available ground-truth datasets, we compare the performance of our metric to widely-used full reference metrics. Results demonstrate that our metric is able to effectively predict the level of distortion in the degraded point cloud contents, achieving high correlation values with respect to subjective scores. Irene Viola 0001, Pablo César |
IEEE Signal Process. Lett. | 2 |
| 2019 | Measuring and Understanding Photo Sharing Experiences in Social Virtual RealityabstractMillions of photos are shared online daily, but the richness of interaction compared with face-to-face (F2F) sharing is still missing. While this may change with social Virtual Reality (socialVR), we still lack tools to measure such immersive and interactive experiences. In this paper, we investigate photo sharing experiences in immersive environments, focusing on socialVR. Running context mapping (N=10), an expert creative session (N=6), and an online experience clustering questionnaire (N=20), we develop and statistically evaluate a questionnaire to measure photo sharing experiences. We then ran a controlled, within-subject study (N=26 pairs) to compare photo sharing under F2F, Skype, and Facebook Spaces. Using interviews, audio analysis, and our questionnaire, we found that socialVR can closely approximate F2F sharing. We contribute empirical findings on the immersiveness differences between digital communication media, and propose a socialVR questionnaire that can in the future generalize beyond photo sharing. Jie Li 0064, Yiping Kong, Thomas Röggla, Francesca De Simone, Swamy Ananthanarayan, Huib de Ridder, Abdallah El Ali, Pablo César |
CHI | 8 |
| 2019 | CorrFeat: Correlation-based Feature Extraction Algorithm using Skin Conductance and Pupil Diameter for Emotion RecognitionabstractTo recognize emotions using less obtrusive wearable sensors, we present a novel emotion recognition method that uses only pupil diameter (PD) and skin conductance (SC). Psychological studies show that these two signals are related to the attention level of humans exposed to visual stimuli. Based on this, we propose a feature extraction algorithm that extract correlation-based features for participants watching the same video clip. To boost performance given limited data, we implement a learning system without a deep architecture to classify arousal and valence. Our method outperforms not only state-of-art approaches, but also widely-used traditional and deep learning methods. Tianyi Zhang 0013, Abdallah El Ali, Chen Wang 0034, Xintong Zhu, Pablo César |
ICMI | 5 |
| 2019 | Watching Videos Together in Social Virtual Reality: An Experimental Study on User's QoEabstractIn this paper, we describe a user study in which pairs of users watch a video trailer and interact with each other, using two social Virtual Reality (sVR) systems, as well as in a face-to-face condition. The sVR systems are: Facebook Spaces, based on puppet-like customized avatars, and a video-based sVR system using photo-realistic virtual user representations. We collect subjective and objective data to analyze users' Quality of Experience (QoE) and compare their interaction in VR to that observed during the real-life scenario. Our results show that the experience delivered by the video-based sVR system is comparable with real-life settings, while the puppet-based avatars limit the perceived quality of the interaction. Our protocol for QoE assessment is fully documented to allow replication in similar experiments. Francesca De Simone, Jie Li 0064, Henrique Debarba, Abdallah El Ali, Simon Gunkel, Pablo César |
VR | 6 |
| 2019 | Introduction to the Best Papers of the ACM Multimedia Systems (MMSys) Conference 2018 and the ACM Workshop on Network and Operating System Support for Digital Audio and Video (NOSSDAV) 2018 and the International Workshop on Mixed and Virtual Environment Systems (MMVE) 2018abstractNo abstract available. Pablo César, Michael Zink, Niall Murray |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2018 | Workflow Support for Live Object-Based BroadcastingabstractThis paper examines the document aspects of object-based broadcasting. Object-based broadcasting augments traditional video and audio broadcast content with additional (temporally-constrained) media objects. The content of these objects -- as well as their temporal validity -- are determined by the broadcast source, but the actual rendering and placement of these objects can be customized to the needs/constraints of the content viewer(s). The use of object-based broadcasting enables a more tailored end-user experience than the one-size-fits-all of traditional broadcasts: the viewer may be able to selectively turn off overlay graphics (such as statistics) during a sports game, or selectively render them on a secondary device. Object-based broadcasting also holds the potential for supporting presentation adaptivity for accessibility or for device heterogeneity. Jack Jansen 0001, Pablo César, Dick C. A. Bulterman |
DocEng | 2 |
| 2018 | 4G/LTE channel quality reference signal trace data setabstractMobile networks, especially LTE networks, are used more and more for high-bandwidth services like multimedia or video streams. The quality of the data connection plays a major role in the perceived quality of a service. Videos may be presented in a low quality or experience a lot of stalling events, when the connection is too slow to buffer the next frames for playback. So far, no publicly available data set exists that has a larger number of LTE network traces and can be used for deeper analysis. In this data set, we provide 546 traces of 5 minutes each with a sample rate of 100 ms. Thereof 377 traces are pure LTE data. We furthermore provide an Android app to gather further traces as well as R scripts to clean, sort, and analyze the data. Britta Meixner, Jan Willem Kleinrouweler, Pablo César |
MMSys | 3 |
| 2018 | Towards Individual QoE for Multiparty VideoconferencingabstractVideoconferencing is becoming an essential part in everyday life. The visual channel allows us for interactions that were not possible over audio-only communication systems, such as the telephone. However, being a de-facto over-the-top service, the quality of the delivered videoconferencing experience is subject to variations, depending on network conditions. Videoconferencing systems adapt to network conditions by changing, for example, encoding bit rate of the video. For this adaptation not to hamper the benefits related to the presence of a video channel in the communication, it needs to be optimized according to a measure of the quality of experience (QoE) as perceived by the user. The latter is highly dependent on the ongoing interaction and individual preferences, which have hardly been investigated so far. In this paper, we focus on the impact that video quality has on conversations that revolve around objects that are presented over the video channel. To this end, we conducted an empirical study where groups of four people collaboratively build a Lego model over a videoconferencing system. We examine the requirements for such a task by showing when the interaction, measured by visual and auditory cues, changes depending on the encoding bit rate and loss. We then explore the impact that prior experience with the technology and affective state have on QoE of participants. We use these factors to construct predictive models that double the accuracy compared to a model based on the system factors alone. We conclude with a discussion of how these factors could be applied in real-world scenarios. Marwin Schmitt, Judith Redi, Dick C. A. Bulterman, Pablo César |
IEEE Trans. Multim. | 4 |
| 2018 | Best Papers of the ACM Multimedia Systems (MMSys) Conference 2017 and the ACM Workshop on Network and Operating System Support for Digital Audio and Video (NOSSDAV) 2017abstractBest Papers of the ACM Multimedia Systems (MMSys) Conference 2017 and the ACM Workshop on Network and Operating System Support for Digital Audio and Pablo César, Cheng-Hsin Hsu, Chun-Ying Huang, Pan Hui 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2017 | Immersion and Togetherness: How Live Visualization of Audience Engagement Can Enhance Music Events
Najereh Shirzadian, Judith Redi, Thomas Röggla, Alice Panza, Frank Nack, Pablo César |
ACE | 6 |
| 2017 | Tangible Air: An Interactive Installation for Visualising Audience EngagementabstractThis article presents an end-to-end system for capturing physiological sensor data and visualising it on a real-time graphic dashboard and as part of an art installation. More specifically, it describes an event where the level of engagement of the audience was measured by means of Galvanic Skin Response (GSR) sensors and of the presenter through a sweater fitted with GSR, ECG and acceleration sensors. The gathered data was presented in real-time through a visualisation projected onto a screen and a physical electro-mechanical installation, which would change the height of helium-filled balloons depending on the atmosphere in the auditorium. Thereby trying to create a tangible way of making the invisible visible. Thomas Röggla, Chen Wang 0034, Lilia Perez Romero, Jack Jansen 0001, Pablo César |
Creativity & Cognition | 5 |
| 2017 | The Play Is a Hit: But How Can You Tell?abstractResearch has shown that physiological sensors provide a valuable mechanism for quantifying the experience of audiences attending cultural events. In the strength of the applause or questionnaires, bio-sensors provide fine-grained timed data that can be used to infer the quality of the experience of the audience members. Unfortunately, available commercial sensors are designed for lab or home usage and studies, focusing on the individual instead of the reactions of a crowd of people. In this study, we present our own designed physiological measurement system that overcomes the challenges of using and deploying sensors in a theatrical environment (anonymity and privacy, real-time gathering of data, support for large crowds of 30-100 people), and the wearable system has two unique features that distinguished between adult audience members and children ones. We report our experimental results which particularly answered the research questions proposed by the producer, director, and the artists. Chen Wang 0034, Pablo César |
Creativity & Cognition | 2 |
| 2017 | Enhancing Music Events Using Physiological Sensor DataabstractThis demo showcases a real-time visualisation displaying the level of engagement of a group of people attending a Jazz concert. Based on wearable sensor technology and machine learning principles, we present how this visualisation for enhancing events was developed following a user-centric approach. We describe the process of running an experiment using our custom physiological sensor platform, gathering requirements for the visualisation and finally implementing said visualisation. The end result being collaborative artwork to enhance people's immersion into cultural events. Thomas Röggla, Najereh Shirzadian, Alice Panza, Pablo César |
ACM Multimedia | 5 |
| 2017 | CWI-ADE2016 Dataset: Sensing nightclubs through 40 million BLE packetsabstractThe CWI-ADE2016 Dataset is a collection of more than 40 million Bluetooth Low Energy (BLE) packets and of 14 million accelerometer and temperature samples generated by wristbands that people wore in a nightclub. The data was gathered during Amsterdam Dance Event 2016 in an exclusive club experience curated around human senses, which leveraged technology as a bridge between the club and the guests. Each guest was handed a custom-made wristband with a BLE-enabled device that broadcast movement, temperature and other sensor readings. A network of Raspberry Pi receivers deployed for the occasion captured broadcast packets from wristbands and any other BLE device in the environment. This data provides a full picture of the performance of the real life deployment of a sensing infrastructure and gives insights to designing sensing platforms, understanding networks and crowds behaviour or studying opportunistic sensing. This paper describes an analysis of this dataset and some examples of usage. Sergio Cabrero, Jack Jansen 0001, Thomas Röggla, John Alexis Guerra Gómez, David A. Shamma, Pablo César |
MMSys | 6 |
| 2017 | Improving Video Quality in Crowded Networks Using a DANEabstractDynamic Adaptive Streaming over HTTP (DASH) is a technology for delivering video content over the Internet. It provides an effective mechanism, which has been adopted by major content providers. Nevertheless, available DASH player implementations have a number of drawbacks such as performance problems on shared network connections, which lead to video freezes and frequent video quality changes. In this paper, we propose a method to reduce the performance problems that exist in networks with a large number of DASH players. These networks can be found in hotels, apartment complexes, and airports. In experiments with up to 600 simultaneously active players, we are able to reduce the number of DASH players with freezes by 95% (from 345 to 15) compared to throughput-based adaptation and by 75% (from 62 to 15) compared to BOLA using our DASH Assisting Network Element (DANE). In addition, we reduced the number of quality switches by 94% compared to throughput-based adaptation, and by 85% compared to BOLA. Jan Willem Kleinrouweler, Britta Meixner, Pablo César |
NOSSDAV | 3 |
| 2017 | Co-present and remote audience experiences: intensity and cohesion
Erik Geelhoed, Kuldip Singh-Barmi, Ian Biscoe, Pablo César, Jack Jansen 0001, Chen Wang 0034, Rene Kaiser |
Multim. Tools Appl. | 4 |
| 2017 | Design, Implementation, and Evaluation of a Point Cloud Codec for Tele-Immersive VideoabstractWe present a generic and real-time time-varying point cloud codec for 3D immersive video. This codec is suitable for mixed reality applications in which 3D point clouds are acquired at a fast rate. In this codec, intra frames are coded progressively in an octree subdivision. To further exploit interframe dependencies, we present an inter-prediction algorithm that partitions the octree voxel space in N × N × N macroblocks (N = 8, 16, 32). The algorithm codes points in these blocks in the predictive frame as a rigid transform applied to the points in the intra-coded frame. The rigid transform is computed using the iterative closest point algorithm and compactly represented in a quaternion quantization scheme. To encode the color attributes, we defined a mapping of color per vertex attributes in the traversed octree to an image grid and use legacy image coding method based on JPEG. As a result, a generic compression framework suitable for realtime 3D tele-immersion is developed. This framework has been optimized to run in real time on commodity hardware for both the encoder and decoder. Objective evaluation shows that a higher rate-distortion performance is achieved compared with available point cloud codecs. A subjective study in a state-of-the-art mixed reality system shows that introduced prediction distortions are negligible compared with the original reconstructed point clouds. In addition, it shows the benefit of reconstructed point cloud video as a representation in the 3D virtual world. The codec is available as open source for integration in immersive and augmented communication applications and serves as a base reference software platform in JTC1/SC29/WG11 (MPEG) for the further development of standardized point-cloud compression solutions. Rufael Mekuria, Kees Blom, Pablo César |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | An SDN Architecture for Privacy-Friendly Network-Assisted DASHabstractDynamic Adaptive Streaming over HTTP (DASH) is the premier technology for Internet video streaming. DASH efficiently uses existing HTTP-based delivery infrastructures implementing adaptive streaming. However, DASH traffic is bursty in nature. This causes performance problems when DASH players share a network connection or in networks with heavy background traffic. The result is unstable and lower quality video. In this article, we present the design and implementation of a so-called DASH Assisting Network Element (DANE). Our system provides target bitrate signaling and dynamic traffic control. These two mechanisms realize proper bandwidth sharing among clients. Our system is privacy friendly and fully supports encrypted video streams. Trying to improve the streaming experience for users who share a network connection, our system increases the video bitrate and reduces the number of quality switches. We show this through evaluations in our Wi-Fi testbed. Jan Willem Kleinrouweler, Sergio Cabrero, Pablo César |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2016 | Distributed Liveness: Understanding How New Technologies Transform Performance ExperiencesabstractWe identify emerging phenomena of distributed liveness, involving new relationships among performers, audiences, and technology. Liveness is a recent, technology-based construct, which refers to experiencing an event in real-time with the possibility for shared social realities. Distributed liveness entails multiple forms of physical, spatial, and social co-presence between performers and audiences across physical and virtual spaces. We interviewed expert performers about how they experience liveness in physically co-present and distributed settings. Findings show that distributed performances and technology need to support flexible social co-presence and new methods for sensing subtle audience responses and conveying engagement abstractly. Andrew M. Webb 0001, Chen Wang 0034, Andruid Kerne, Pablo César |
CSCW | 4 |
| 2016 | MP3DG-PCC, Open Source Software Framework for Implementation and Evaluation of Point Cloud CompressionabstractWe present MP3DG-PCC, an open source framework for design, implementation and evaluation of point cloud compression algorithms. The framework includes objective quality metrics, lossy and lossless anchor codecs, and a test bench for consistent comparative evaluation. The framework and proposed methodology is in use for the development of an international point cloud compression standard in MPEG. In addition, the library is integrated with the popular point cloud library, making a large number of point cloud processing available and aligning the work with the broader open source community. Rufael Mekuria, Pablo César |
ACM Multimedia | 2 |
| 2016 | Delivering stable high-quality video: an SDN architecture with DASH assisting network elementsabstractDynamic adaptive streaming over HTTP (DASH) is a simple, but effective, technology for video streaming over the Internet. It provides adaptive streaming while being highly scalable at the side of the content providers. However, the mismatch between TCP and the adaptive bursty nature of DASH traffic results in underperformance of DASH streams in busy networks. This paper describes a networking architecture based on the Software Defined Networking (SDN) paradigm. Controllers in the network with a broad overview on the network activity provide two mechanisms for adaptation assistance: explicitly signaling target bitrates to DASH players and dynamic traffic control in the network. We evaluate how each of these mechanisms can contribute to the delivery of a stable and high quality stream. It shows that our architecture improves the quality of experience by doubling the video bitrate and reducing disturbing quality switches. As such, this paper contributes insights on how to implement DASH-aware networking that also enables internet service providers, network administrators, and end-users to configure their networks to their requirements. Jan Willem Kleinrouweler, Sergio Cabrero, Pablo César |
MMSys | 3 |
| 2016 | 1Mbps is enough: Video quality and individual idiosyncrasies in multiparty HD video-conferencingabstractMost video platforms deliver HD video in high bitrate encoding. Modern video-conferencing systems are capable of handling HD streams, but using multiparty conferencing, average internet connections in the home are on their bandwidth limit. For properly managing the encoding bitrate in videoconferencing, we must know what is the minimum bitrate requirement to provide users an acceptable experience, and what is the bitrate level after which QoE saturates?. Most available subjective studies in this area used rather dated technologies. We report on a multiparty study on video quality with HD resolution. We tested different encoding bitrates (256kbs, 1024kbs and 4096kbs) and packet loss rates (0, 0.5%) in groups of 4 participants with a scenario based on the ITU building blocks task. We discuss the influence of group interaction and individual idiosyncrasies based on different mixed models, and look at covariates engagement and enjoyment as further explanatory factors. We found that 256kbs is still sufficient to provide a fair overall experience, but video quality is noticed to be poor. On the higher bitrate end, most people will not perceive the difference between 1024kbs and 4096kbs, considering in both cases the quality to be close to excellent. Independent on bitrate, packet loss has a small but significant impact, quantifiable in, on average, less than half a point difference on a 5-point ITU scale. Marwin Schmitt, Judith Redi, Pablo César, Dick C. A. Bulterman |
QoMEX | 3 |
| 2016 | Quantifying audience experience in the wild: Heuristics for developing and deploying a biosensor infrastructure in theatersabstractMeasuring the experience of audience of arts events is essential in the “experience economy” of this day and age, but it is a difficult task. The value of such information goes beyond evaluating the impact of the arts, as it can provide insights and feedback to enhance the work of artists and the experiences of other audience members. Through in-depth understanding of the needs of the providers and consumers of the arts, we progressively developed a biosensor infrastructure that was deployed in theaters. Over the years, we identified the challenges and issues related to developing and deploying a biosensor infrastructure in theaters. These collective experiences and identified issues were categorized into three main areas: processes, data, and system. A total of seven heuristics are developed across the three main areas. Processes place the stakeholders and audiences at the core of the research; data provides guidelines for data validity, collecting a variety of data, and supporting real-time data gathering; and systems covers the concurrency, scalability, deployment and feedback of the infrastructure. We believe that this set of heuristics forms the foundation for an adequate infrastructure to measure audience experience in the wild and it is a valuable source of guideline for future work. Chen Wang 0034, Jacqueline Wong, Xintong Zhu, Thomas Röggla, Jack Jansen 0001, Pablo César |
QoMEX | 6 |
| 2016 | A model for evaluating sharing policies for network-assisted HTTP adaptive streaming
Jan Willem Kleinrouweler, Sergio Cabrero, Robert D. van der Mei, Pablo César |
Comput. Networks | 4 |
| 2015 | Multimedia Document Structure for Distributed TheatreabstractThis paper explores the suitability of structured (and declarative) multimedia document formats for supporting a novel type of performing arts: distributed theatre. In distributed theatre, the actors are split between two (or more) locations, but together deliver a single performance mediated by the cameras, the internet, and projection technologies. Based on our efforts to make an actual distributed theatre production happen (the Tempest by Miracle Theatre), this paper reflects on our experience. Our findings are divided into two main areas: workflow and document structure. We conclude that novel types of video-mediated applications, like distributed theatre, require new manners of authoring documents. Moreover, specific extensions to existing document formats are needed in order to accommodate the new requirements imposed by such kind of applications. Jack Jansen 0001, Michael Frantzis, Pablo César |
DocEng | 3 |
| 2015 | Analysing Audience Response to Performing Events: A Web Platform for Interactive Exploration of Physiological Sensor DataabstractThis paper presents a web interface for the exploration of audience response to a performing arts event. The platform temporally synchronises data obtained from physiological sensors with video recordings. More concretely, it is geared towards people from the creative industry, e.g. theatre directors, who want to gain deeper insights into how audiences perceive their performances. The platform takes the raw data from sensors and corresponding video recordings and visualises both synchronised in a more digestible manner. This document presents the major features of the application and explains the reasoning behind its various visualisations and interactive capabilities and how it can benefit performing artists. Thomas Röggla, Chen Wang 0034, Pablo César |
ACM Multimedia | 3 |
| 2015 | A Distributed Theatre Experiment with ShakespeareabstractThis paper reports on an experimental production of The Tempest that was developed in collaboration with Miracle Theatre Company realised as a distributed performance from two separate stages through a dynamically configured telepresence system. The production allowed an exploration of the way a range of technologies, including consumer grade broadband, cameras and projection technologies could affect the development and delivery of live theatre by regional touring company. The architecture of the communication platform used to deliver the performance is introduced as are two novel software tools that are used to describe and control the way the play should be captured and represented. Doug Williams, Ian Kegel, Marian Florin Ursu, Pablo César, Jack Jansen 0001, Erik Geelhoed, Andras Horti, Michael Frantzis, Bill Scott |
ACM Multimedia | 4 |
| 2015 | Forward to the theme issue on interactive experiences for television and online video
Marianna Obrist, Pablo César, Santosh Basapur |
Pers. Ubiquitous Comput. | 2 |
| 2014 | Sensing a live audienceabstractPsychophysiological measurement has the potential to play an important role in audience research. Currently, such research is still in its infancy and it usually involves collecting data in the laboratory, where during each experimental session one individual watches a video recording of a performance. We extend the experimental paradigm by simultaneously measuring Galvanic Skin Response (GSR) of a group of participants during a live performance. GSR data were synchronized with video footage of performers and audience. In conjunction with questionnaire data, this enabled us to identify a strongly correlated main group of participants, describe the nature of their theatre experience and map out a minute-by-minute unfolding of the performance in terms of psycho-physiological engagement. The benefits of our approach are twofold. It provides a robust and accurate mechanism for assessing a performance. Moreover, our infrastructure can enable, in the future, real-time feedback from remote audiences for online performances. Chen Wang 0034, Erik Geelhoed, Phil Stenton, Pablo César |
CHI | 4 |
| 2014 | Low complexity connectivity driven dynamic geometry compression for 3D Tele-ImmersionabstractGeometry based 3D Tele-Immersion is a novel emerging media application that involves on the fly reconstructed 3D mesh geometry. To enable real-time communication of such live reconstructed mesh geometry over a bandwidth limited link, fast dynamic geometry compression is needed. However, most tools and methods have been developed for compressing synthetically generated graphics content. These methods achieve good compression rates by exploiting topological and geometric properties that typically do not hold for reconstructed mesh geometry. The live reconstructed dynamic geometry is causal and often non-manifold, open, non-oriented and time-inconsistent. Based on our experience developing a prototype for 3D Teleimmersion based on live reconstructed geometry, we discuss currently available tools. We then present our approach for dynamic compression that better exploits the fact that the 3D geometry is reconstructed and achieve a state of art rate-distortion under stringent real-time constraints. Rufael Mekuria, Pablo César, Dick C. A. Bulterman |
ICASSP | 2 |
| 2014 | 3rd International Workshop on Socially-Aware Multimedia (SAM'14)abstractMultimedia social communication is becoming commonplace. Television is becoming smart and social; media sharing applications are transforming the way we converse and recall events and videoconferencing is a common application on our computers, phones, tablets and even televisions. The confluence of computer-mediated interaction, social networking, and multimedia content are radically reshaping social communications, bringing new challenges and opportunities. This workshop, in its third edition, provides an opportunity to explore socially-aware multimedia, in which the social dimension of mediated interactions between people are considered to be as important as the characteristics of the media content. Even though this social dimension is implicitly addressed in some current solutions, further research is needed to better understand what makes multimedia socially-aware. Pablo César, David A. Shamma, Matthew Cooper 0002, Aisling Kelliher |
ACM Multimedia | 1 |
| 2014 | Design, development and assessment of control schemes for IDMS in a standardized RTCP-based solution
Mario Montagud, Fernando Boronat, Hans Stokking, Pablo César |
Comput. Networks | 4 |
| 2014 | Multimedia authoring and annotation
Dick C. A. Bulterman, Pablo César, Ethan V. Munson, Maria da Graça Campos Pimentel |
Multim. Tools Appl. | 2 |
| 2014 | Enabling Geometry-Based 3-D Tele-Immersion With Fast Mesh Compression and Linear Rateless Codingabstract3-D tele-immersion (3DTI) enables participants in remote locations to share, in real time, an activity. It offers users interactive and immersive experiences, but it challenges current media-streaming solutions. Work in the past has mainly focused on the efficient delivery of image-based 3-D videos and on realistic rendering and reconstruction of geometry-based 3-D objects. The contribution of this paper is a real-time streaming component for 3DTI with dynamic reconstructed geometry. This component includes both a novel fast compression method and a rateless packet protection scheme specifically designed towards the requirements imposed by real time transmission of live-reconstructed mesh geometry. Tests on a large dataset show an encoding speed-up up to ten times at comparable compression ratio and quality, when compared with the high-end MPEG-4 SC3DMC mesh encoders. The implemented rateless code ensures complete packet loss protection of the triangle mesh object and a delivery delay within interactive bounds. Contrary to most linear fountain codes, the designed codec enables real-time progressive decoding allowing partial decoding each time a packet is received. This approach is compared with transmission over TCP in packet loss rates and latencies, typical in managed WAN and MAN networks, and heavily outperforms it in terms of end-to-end delay. The streaming component has been integrated into a larger 3DTI environment that includes state of the art 3-D reconstruction and rendering modules. This resulted in a prototype that can capture, compress transmit, and render triangle mesh geometry in real-time in realistic internet conditions as shown in experiments. Compared with alternative methods, lower interactive end-to-end delay and frame rates over three times higher are achieved. Rufael Mekuria, Michele Sanna, Ebroul Izquierdo, Dick C. A. Bulterman, Pablo César |
IEEE Trans. Multim. | 5 |
| 2013 | Multimedia document synchronization in a distributed social contextabstractWatching digital content together and commenting on it is becoming a social habit between friends and family members living apart. It is also becoming an important value-added activity for business video conferencing. In both cases, the video sharing experience can easily be spoiled if synchronization problems arise, since the context of the conversation will not be consistent across locations. In the past, research has treated the distributed synchronization problem as a technical one, mainly focusing on timestamps, frame accuracy, and protocol-dependent control messages. That approach is based on a content agnostic approach which we feel does not adequately address the higher-level constraints of individual conversations. Jack Jansen 0001, Pablo César, Dick C. A. Bulterman |
ACM Symposium on Document Engineering | 2 |
| 2013 | A Quality of Experience Testbed for Video-Mediated Group CommunicationabstractVideo-Mediated group communication is quickly moving from the office to the home, where network conditions might fluctuate. If we are to provide a software component that can, in real-time, monitor the Quality of Experience (QoE), we would have to carry out extensive experiments under different varying (but controllable) conditions. Unfortunately, there are no tools available that provide us the required fined-grained level of control. This paper reports on our efforts implementing such a test bed. The test bed provides the experiment conductor full control over the complete media pipeline, and the possibility of modifying in real-time network and media conditions. Additionally, it has facilities to easily develop an experiment with custom layouts, task integration, and assessment of subjective ratings through questionnaires. We have already used the test bed in a number of evaluations, reported in this paper for discussing the benefits and drawbacks of our solution. The test bed have been proven to be a flexible and effective canvas for better understanding QoE on video-mediated group communication. Marwin Schmitt, Simon Gunkel, Pablo César |
ISM | 3 |
| 2013 | 2nd international workshop on socially-aware multimedia (SAM'13)abstractMultimedia social communication is becoming commonplace. Television is becoming smart and social; media sharing applications are transforming the way we converse and recall events and videoconferencing is a common application on our computers, phones, tablets and even televisions. The confluence of computer-mediated interaction, social networking, and multimedia content are radically reshaping social communications, bringing new challenges and opportunities. This workshop, in its second edition, provides an opportunity to explore socially-aware multimedia, in which the social dimension of mediated interactions between people are considered to be as important as the characteristics of the media content. Even though this social dimension is implicitly addressed in some current solutions, further research is needed to better understand what makes multimedia socially-aware. Pablo César, Matthew Cooper 0002, David A. Shamma, Doug Williams |
ACM Multimedia | 1 |
| 2013 | A 3D tele-immersion system based on live captured mesh geometryabstract3D Tele-immersion enables participants in remote locations to share, in real-time, an activity. It offers users natural interactivity and immersive experiences, but it challenges current networking solutions. Work in the past has mainly focused on the efficient delivery of image-based 3D videos and on the realistic rendering and reconstruction of geometry-based 3D objects. The contribution of this paper is a complete media pipeline that allows for geometry-based 3D tele-immersion. Unlike previous approaches, that stream videos or video plus depth estimate, our streaming module can transmit the live-reconstructed 3D representations (triangle meshes). Based on a set of comparative experiments, this paper details the architecture and describes a novel component that can efficiently stream geometry in real-time. This component includes both a novel fast local compression algorithm and a rateless packet protection scheme geared towards the requirements imposed by real-time transmission of live-capture mesh geometry. Tests on a large dataset show an encoding and decoding speed-up of over 10 times at similar compression and quality rates, when compared to the high-end MPEG-4 SC3DMC mesh encoder. The implemented rateless code ensures complete packet loss protection of the triangle mesh object and avoids delay introduced by retransmissions. This approach is compared to a streaming mechanism over TCP and outperforms it at packet loss rates over 2% and/or latencies over 9 ms in terms of end-to-end transmission delay. As reported in this paper, the component has been successfully integrated into a larger tele-immersive environment that includes beyond state of the art 3D reconstruction and rendering modules. This resulted in a prototype that can capture, compress transmit and render triangle mesh geometry in real-time over the internet. Rufael Mekuria, Michele Sanna, Stefano Asioli, Ebroul Izquierdo, Dick C. A. Bulterman, Pablo César |
MMSys | 6 |
| 2013 | Socially-aware multimedia authoring: Past, present, and futureabstractCreating compelling multimedia productions is a nontrivial task. This is as true for creating professional content as it is for nonprofessional editors. During the past 20 years, authoring networked content has been a part of the research agenda of the multimedia community. Unfortunately, authoring has been seen as an initial enterprise that occurs before ‘real’ content processing takes place. This limits the options open to authors and to viewers of rich multimedia content for creating and receiving focused, highly personal media presentations. This article reflects on the history of multimedia authoring. We focus on the particular task of supporting socially-aware multimedia , in which the relationships within particular social groups among authors and viewers can be exploited to create highly personal media experiences. We provide an overview of the requirements and characteristics of socially-aware multimedia authoring within the context of exploiting community content. We continue with a short historical perspective on authoring support for these types of situations. We then present an overview of a current system for supporting socially-aware multimedia authoring within the community content. We conclude with a discussion of the issues that we feel can provide a fruitful basis for future multimedia authoring support. We argue that providing support for socially-aware multimedia authoring can have a profound impact on the nature and architecture of the entire multimedia information processing pipeline. Dick C. A. Bulterman, Pablo César, Rodrigo Laiola Guimarães |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2012 | Just-in-time personalized video presentationsabstractUsing high-quality video cameras on mobile devices, it is relatively easy to capture a significant volume of video content for community events such as local concerts or sporting events. A more difficult problem is selecting and sequencing individual media fragments that meet the personal interests of a viewer of such content. In this paper, we consider an infrastructure that supports the just-in-time delivery of personalized content. Based on user profiles and interests, tailored video mash-ups can be created at view-time and then further tailored to user interests via simple end-user interaction. Unlike other mash-up research, our system focuses on client-side compilation based on personal (rather than aggregate) interests. This paper concentrates on a discussion of language and infrastructure issues required to support just-in-time video composition and delivery. Using a high school concert as an example, we provide a set of requirements for dynamic content delivery. We then provide an architecture and infrastructure that meets these requirements. We conclude with a technical and user analysis of the just-in-time personalized video approach. Jack Jansen 0001, Pablo César, Rodrigo Laiola Guimarães, Dick C. A. Bulterman |
ACM Symposium on Document Engineering | 2 |
| 2012 | International workshop on socially-aware multimedia (SAM'12)abstractMultimedia social communication is filtering into everyday use. Videoconferencing is appearing in the living room and beyond, television is becoming smart and social, and media sharing applications are transforming the way we converse and recall events. The confluence of computer-mediated interaction, social networking, and multimedia content are radically reshaping social communications, bringing new challenges and opportunities. This workshop provides an opportunity to explore socially-aware multimedia, in which the social dimension of mediated interactions between people are considered as important as the characteristics of the media content. Even though this social dimension is implicitly addressed in some current solutions, further research is needed to better understand what makes multimedia socially-aware. In other words, social interactivity needs to become a first class citizen of multimedia research. Pablo César, David A. Shamma, Doug Williams, Cees Snoek |
ACM Multimedia | 1 |
| 2012 | Enabling 'togetherness' in high-quality domestic videoabstractLow-cost video conferencing systems have provided an existence proof for the value of video communication in a home setting. At the same time, current systems have a number of fundamental limitations that inhibit more general social interactions among multiple groups of participants. In our work, we describe the development, implementation and evaluation of a domestic video conferencing system that is geared to providing true 'togetherness' among conference participants. We show that such interactions require sophisticated support for high-quality audiovisual presentation, and processing support for person identification and localisation. In this paper, we describe user requirements for effective interpersonal interaction. We then report on a system that implements these requirements. We conclude with a systems and user evaluation of this work. We present results that show that participants in a video conference can be made feel as 'together' as collocated players of a board game. Ian Kegel, Pablo César, Jack Jansen 0001, Dick C. A. Bulterman, Tim Stevens, Joke Kort, Nikolaus Färber |
ACM Multimedia | 2 |
| 2012 | Introduction to the Special Section on Smart, Social, and Converged TVabstractThe seven papers in this special section focus om recent advances in the growing research field of television. Oscar Martínez Bonastre, Marie-José Montpetit, Pablo César, Jon Crowcroft, Maja Matijasevic, Zhu Liu 0001 |
IEEE Trans. Multim. | 3 |
| 2011 | Past, present, and future of social TV: A categorizationabstractSocial Television constitutes a fundamental shift in how people interact and socialize around television content. Websites are starting to combine streaming services with social networking sites such as Facebook and Twitter. Media software like Boxee allows users to recommend and share favorite television programs, and friends can jointly watch television remotely as demonstrated by Motorola's Social TV. The purpose of this paper is to provide a structured overview of current developments in this emerging field. Current offerings can be categorized based on the social purpose: content selection and recommendation, communication, community building, and status update. The framework proposed in this article is useful for better understanding the current situation and for identifying future developments. Pablo César, David Geerts |
CCNC | 1 |
| 2011 | Are we in sync?: synchronization requirements for watching online video togetherabstractSynchronization between locations is an important factor for enabling remote shared experiences. Still, experimental data on what is the acceptable synchronization level is scarce. This paper discusses the synchronization requirements for watching online videos together - a popular set of services that recreate the shared experience of watching TV together by offering tools to communicate while watching. It studies the noticeability and annoyance of synchronization differences of the video being watched, as well as the impact on users' feelings of togetherness, both for voice chat and text chat. Results of an experiment with 36 participants show that when using voice chat, users notice synchronization differences sooner, are more annoyed and feel more together than when using text chat. However, users with high text chat activity notice synchronization differences similar to participants using voice chat. David Geerts, Ishan Vaishnavi, Rufael Mekuria, M. Oskar van Deventer, Pablo César |
CHI | 5 |
| 2011 | Multimedia document processing in an HTML5 worldabstractThe evolution in media support within W3C standards has led to the development of HMTL5. HTML5 provides extensive support for audio/video/timed-text within an interoperable browser context. Dick C. A. Bulterman, Rodrigo Laiola Guimarães, Pablo César, Ethan V. Munson, Maria da Graça Campos Pimentel |
ACM Symposium on Document Engineering | 3 |
| 2011 | Towards synergy between the open source and the research multimedia communitiesabstractThis panel extends current efforts from the ACM Multimedia 2011 Organization Committee in taking an important step towards open source projects. The panelists include speakers who are among the leading figures from the open source community. The goal is to provide a shared space for discussion and interaction among consolidated and new open source projects and multimedia researchers. Pablo César, Wei Tsang Ooi, Ben Moskowitz, Zohar Babin, Dick C. A. Bulterman, Rainer Lienhart, Robert Richter |
ACM Multimedia | 1 |
| 2011 | Creating personalized memories from social events: community-based support for multi-camera recordings of school concertsabstractThe wide availability of relatively high-quality cameras makes it easy for many users to capture video fragments of social events such as concerts, sports events or community gatherings. The wide availability of simple sharing tools makes it nearly as easy to upload individual fragments to on-line video sites. Current work on video mashups focuses on the creation of a video summary based on the characteristics of individual media fragments, but it fails to address the interpersonal relationships and time-variant social context within a community of diverse (but related) users. The aim of this paper is to reformulate the research problem of video authoring, by investigating the social relationships of the media 'authors' relative to the performers. Based on a 10-month evaluation process, we specify a set of guidelines for the design and implementation of socially-aware video editing and sharing tools. Our contributions have been realized and evaluated in a prototype software that enables community-based users to navigate through a large common content space and to generate highly personalized video compilations of targeted interest within a social circle. According to the results, a system like ours is a valid alternative for social interactions when apart. We hope that our insights can stimulate future research on socially-aware multimedia tools and applications. Rodrigo Laiola Guimarães, Pablo César, Dick C. A. Bulterman, Vilmos Zsombori, Ian Kegel |
ACM Multimedia | 2 |
| 2011 | Accurate and low-delay seeking within and across mash-ups of highly-compressed videosabstractIn typical video mash-up systems, a group of source videos are compiled off-line into a single composite object. This improves rendering performance, but limits the possibilities for dynamic composition of personalized content. This paper discusses systems and network issues for enabling client-side dynamic composition of video mash-ups. In particular, this paper describes a novel algorithm to support accurate, low-delay seamless composition of independent clips. We report on an intelligent application-steered scheme that allows system layers to prefetch and discard predicted frames before the rendering moment of indexed content. This approach unifies application-level quality-of-experience specification with system layer quality-of-service processing. To evaluate our scheme, several experiments are conducted and substantial performance improvements are observed in terms of accuracy and low delay. Jack Jansen 0001, Pablo César, Dick C. A. Bulterman |
NOSSDAV | 3 |
| 2011 | Guest editorial: Networked television
Cristian Hesselman, Pablo César, David Geerts |
Multim. Syst. | 2 |
| 2011 | IPTV: challenges and future directions
Oscar Martínez Bonastre, Marie-José Montpetit, Pablo César |
Multim. Tools Appl. | 3 |
| 2011 | Advances in IPTV technologies
Oscar Martínez Bonastre, Marie-José Montpetit, Pablo César |
Signal Process. Image Commun. | 3 |
| 2011 | From IPTV to synchronous shared experiences challenges in design: Distributed media synchronization
Ishan Vaishnavi, Pablo César, Dick C. A. Bulterman, Oliver Friedrich, Simon Gunkel, David Geerts |
Signal Process. Image Commun. | 2 |
| 2011 | Enabling Composition-Based Video-Conferencing for the HomeabstractThis paper describes a videoconferencing system that meets performance constraints and functional requirements for use in consumer homes. Our system improves existing home technologies (such as video chat) by providing high-quality audiovisual communication, efficient encoding mechanisms, and low end-to-end delay. Moreover, the system includes a control interface that is capable of dynamically manipulating and compositing audiovisual content streams. This innovative architectural component is required for a domestic setting, where the television acts as the main screen and multiple people gather around it. Apart from the requirements and architecture, this paper analyses the performance of our system. The results validate our architectural decisions and provide a valuable input for further research in domestic videoconferencing. Jack Jansen 0001, Pablo César, Dick C. A. Bulterman, Tim Stevens, Ian Kegel, J. Issing |
IEEE Trans. Multim. | 2 |
| 2010 | Creating and sharing personalized time-based annotations of videos on the webabstractThis paper introduces a multimedia document model that can structure community comments about media. In particular, we describe a set of temporal transformations for multimedia documents that allow end-users to create and share personalized timed-text comments on third party videos. The benefit over current approaches lays in the usage of a rich captioning format that is not embedded into a specific video encoding format. Using as example a Web-based video annotation tool, this paper describes the possibility of merging video clips from different video providers into a logical unit to be captioned, and tailoring the annotations to specific friends or family members. In addition, the described transformations allow for selective viewing and navigation through temporal links, based on end-users' comments. We also report on a predictive timing model for synchronizing unstructured comments with specific events within a video(s). The contributions described in this paper bring significant implications to be considered in the analysis of rich media social networking sites and the design of next generation video annotation tools. Rodrigo Laiola Guimarães, Pablo César, Dick C. A. Bulterman |
ACM Symposium on Document Engineering | 2 |
| 2010 | A model for editing operations on active temporal multimedia documentsabstractInclusion of content with temporal behavior in a structured document leads to such a document gaining temporal semantics. If we then allow changes to the document during its presentation, this brings with it a number of fundamental issues that are related to those temporal semantics. In this paper we study modifications of active multimedia documents and the implications of those modifications for temporal consistency. Such modifications are becoming increasingly important as multimedia documents move from being primarily a standalone presentation format to being a building block in a larger application. Jack Jansen 0001, Pablo César, Dick C. A. Bulterman |
ACM Symposium on Document Engineering | 2 |
| 2010 | From IPTV services to shared experiences: Challenges in architecture designabstractThis paper discusses the architectural challenges of transitioning from services to experiences. In particular, it focuses on evolution from traditional IPTV services to more social scenarios, in which groups of people in different locations watch synchronized multimedia content together. In addition to the multimedia content, the shared experiences envisioned in this article provide a real-time communication channel between the participants. Based on an implemented architecture, this paper identifies a number of challenges and analyze them. The most important challenges highlighted in this article include: shared experience modeling, universal session handling, synchronization, and quality of service. This article is the first stone paving the way for a truly interoperable ecosystem, which can offer cross-domain experiences to the users. Ishan Vaishnavi, Pablo César, Dick C. A. Bulterman, Oliver Friedrich |
ICME | 2 |
| 2009 | Adding dynamic visual manipulations to declarative multimedia documentsabstractThe objective of this work is to define a document model extension that enables complex spatial and temporal interactions within multimedia documents. As an example we describe an authoring interface of a photo sharing system that can be used to capture stories in an open, declarative format. The document model extension defines visual transformations for synchronized navigation driven by dynamic associated content. Due to the open declarative format, the presentation content can be targeted to individuals, while maintaining the underlying data model. The impact of this work is reflected in its recent standardization in the W3C SMIL language. Multimedia players, as Ambulant and the RealPlayer, support the extension described in this paper. Fons Kuijk, Rodrigo Laiola Guimarães, Pablo César, Dick C. A. Bulterman |
ACM Symposium on Document Engineering | 3 |
| 2009 | Leveraging user impact: an architecture for secondary screens usage in interactive television
Pablo César, Dick C. A. Bulterman, Jack Jansen 0001 |
Multim. Syst. | 1 |
| 2009 | Fragment, tag, enrich, and send: Enhancing social sharing of videoabstractThe migration of media consumption to personal computers retains distributed social viewing, but only via nonsocial, strictly personal interfaces. This article presents an architecture, and implementation for media sharing that allows for enhanced social interactions among users. Using a mixed-device model, our work allows targeted, personalized enrichment of content. All recipients see common content, while differentiated content is delivered to individuals via their personal secondary screens. We describe the goals, architecture, and implementation of our system in this article. In order to validate our results, we also present results from two user studies involving disjoint sets of test participants. Pablo César, Dick C. A. Bulterman, Jack Jansen 0001, David Geerts, Hendrik Knoche, William Seager |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2008 | Multimedia content transformation: fragmentation, enrichment, and adaptationabstractThis working session will be an interactive discussion about multimedia content transformation. The basic assumption is that content transformation activities should be provided as non-destructive operations. The final goal of the panel is to gather researchers within the community interested in manipulating multimedia content for providing rich user experiences. The organizers of the panel will moderate and shape the discussion; nevertheless, position papers from the participants are expected. Pablo César, Dick C. A. Bulterman, Jack Jansen 0001, Maria da Graça Campos Pimentel, Simone D. J. Barbosa |
ACM Symposium on Document Engineering | 1 |
| 2008 | Multimedia adaptation in ubiquitous environments: benefits of structured multimedia documentsabstractThis paper demonstrates the advantages of using structured multimedia documents for session management and media distribution in ubiquitous environments. We show how document manipulations can be used to perform powerful operations such as content to context adaptation and presentation continuity. When consuming media in ubiquitous environments, where the set of devices surrounding a user may change, dynamic media adaptation and session transfer become primary requirements. This paper presents a working system, based on a representative scenario, in which multimedia content is distributed and adapted to a movable user to best suit his/her contextual situation. The implemented scenario includes the following scenes: content selection using a personal mobile phone, content distribution to the most suitable device according to the user's context, and presentation continuity when the user moves to another location. This paper introduces the underlying document manipulations that turn the scenario into a working system. Pablo César, Ishan Vaishnavi, Ralf Kernchen, Stefan Meissner, Cristian Hesselman, Matthieu Boussard, Antonietta Spedalieri, Dick C. A. Bulterman |
ACM Symposium on Document Engineering | 1 |
| 2008 | Enhancing social sharing of videos: fragment, annotate, enrich, and shareabstractMedia consumption is an inherently social activity, serving to communicate ideas and emotions across both small- and large-scale communities. The migration of the media experience to personal computers retains social viewing, but typically only via a non-social, strictly personal interface. This paper presents an architecture and implementation for media content selection, content (re)organization, and content sharing within a user community that is heterogeneous in terms of both participants and devices. In addition, our application allows the user to enrich the content as a differentiated personalization activity targeted to his/her peer-group. We describe the goals, architecture and implementation of our system in this paper. In order to validate our results, we also present results from two user studies involving disjoint sets of test participants. Pablo César, Dick C. A. Bulterman, David Geerts, Jack Jansen 0001, Hendrik Knoche, William Seager |
ACM Multimedia | 1 |
| 2008 | A presentation layer mechanism for multimedia playback mobility in service oriented architecturesabstractThis paper presents an approach for media presentation continuity in playback mode. We use the term presentation continuity over session transfer since our solution is at the presentation layer. Previous research on this topic has focused on transferring a particular stream or set of related streams at the sessions layer for live broadcasting or conferencing sessions. We argue that in the realm of service oriented architectures, such as telecom operator networks, this approach does not take full advantage of the particular case of media playback. Our mechanism presents an alternative to the traditional approach, which i) Lowers network control plane overhead, thus reducing chances of presentation consistency loss ii) Lowers network data overhead due to lesser need for transcoding iii) Delegates presentation consistency issues, such as inter-media synchronisation, to the media player iv) Dynamically adapts the presentation to the new target devices without transcoding. Finally, we present experimental results which show that our approach is implementable in acceptable time bounds. Ishan Vaishnavi, Pablo César, Jack Jansen 0001, Dick C. A. Bulterman |
MUM | 2 |
| 2008 | A mechanism for presentation-layer media continuity in media playback modeabstractThis demo presents a new approach for media presentation continuity in playback mode. We use the term presentation continuity over session transfer since our solution is at the presentation layer. Previous research on this topic has focused on transferring a particular stream or set of related streams at the sessions layer. Our approach presents an alternative, recognising the fact that a user is connected to a media presentation, which, may be composed of multiple sessions. The advantages of our approach are i) Lower network control plane overhead, thus reducing chances of semantic presentation loss ii) Lower network data overhead due to lesser need for transcoding iii) delegating presentation semantic issues, such as inter-media synchronisation, to the player iv) dynamically adapt the presentation to the new target devices without transcoding. Ishan Vaishnavi, Pablo César, Jack Jansen 0001, Dick C. A. Bulterman |
NOSSDAV | 2 |
| 2008 | Multimedia systems, languages, and infrastructures for interactive television
Pablo César, Dick C. A. Bulterman, Konstantinos Chorianopoulos, Jens F. Jensen |
Multim. Syst. | 1 |
| 2008 | Introduction to special issue: Human-centered television - directions in interactive digital television researchabstractThe research area of interactive digital TV is in the midst of a significant revival. Unlike the first generation of digital TV, which focused on producer concerns that effectively limited (re)distribution, the current generation of research is closely linked to the role of the user in selecting, producing, and distributing content. The research field of interactive digital television is being transformed into a study of human-centered television. Our guest editorial reviews relevant aspects of this transformation in the three main stages of the content lifecycle: content production, content delivery, and content consumption. While past research on content production tools focused on full-fledged authoring tools for professional editors, current research studies lightweight, often informal end-user authoring systems. In terms of content delivery, user-oriented infrastructures such as peer-to-peer are being seen as alternatives to more traditional broadcast solutions. Moreover, end-user interaction is no longer limited to content selection, but now facilitates nonlinear participatory television productions. Finally, user-to-user communication technologies have allowed television to become a central component of an interconnected social experience. The background context given in this article provides a framework for appreciating the significance of four detailed contributions that highlight important directions in transforming interactive television research. Pablo César, Dick C. A. Bulterman, Luiz F. G. Soares |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2007 | An efficient, streamable text format for multimedia captions and subtitlesabstractIn spite of the high profile of media types such as video, audio and images, many multimedia presentations rely extensively on text content. Text can be used for incidental labels, or as subtitles or captions that accompany other media objects. In a multimedia document, text content is not only constrained by the need to support presentation styles and layout, it is also constrained by the temporal context of the presentation. This involves intra-text and extra text timing synchronization with other media objects. This paper describes a new timed-text representation language that is intended to be embedded in a non-text host language. Our format, which we call aText (for the Ambulant Text Format), balances the need for text styling with the requirement for an efficient representation that can be easily parsed and scheduled at runtime. aText, which can also be streamed, is defined as an embeddable text format for use within declarative XML languages. The paper presents a discussion of the requirements for the format, a description of the format and a comparison with other existing and emerging text formats. We also provide examples for aText when embedded within the SMIL and MLIF languages and discuss our implementation experiences of aText with the Ambulant Player. Dick C. A. Bulterman, Jack Jansen 0001, Pablo César, Samuel Cruz-Lara |
ACM Symposium on Document Engineering | 3 |
| 2006 | Benefits of structured multimedia documents in IDTV: the end-user enrichment systemabstractThis paper presents a system that exploits the benefits of modelling multimedia presentations as structured documents within the context of interactive digital television systems. Our work permits end-users to easily enrich multimedia content at viewing time (e.g., add images and delete scenes). Because the document is structured, the system can expose to the user the possible enrichment alternatives depending on the current state of the presentation (e.g., current story). Moreover, because the base content is wrapped as a structured document, the enrichments can be modelled as overlying layers that do not alter the original content. Finally, the user can share the enriched content (or parts of it) to specific peers within a P2P network. Pablo César, Dick C. A. Bulterman, Jack Jansen 0001 |
ACM Symposium on Document Engineering | 1 |
| 2006 | The ambulant annotator: empowering viewer-side enrichment of multimedia contentabstractThis paper presents a set of demos that allow viewer-side enrichment of multimedia content in a home setting. The most relevant features of our system are the following: passive authoring of content in contraposition to the traditional active PC authoring, preservation of the base content, and collaborative authoring (e.g., to share the enriched material with a peer group). These requirements are met by modelling television content as structured multimedia documents using SMIL 2.1. Pablo César, Dick C. A. Bulterman, Jack Jansen 0001 |
ACM Symposium on Document Engineering | 1 |
| 2006 | An architecture for viewer-side enrichment of TV contentabstractThis paper presents a user interface model and implementation for exploiting next-generation interactive capabilities with the domain of television content. Our work studies capabilities that extend a user's potential impact over the consumption and sharing of television programs. The main capabilities of our environment include personalized viewing and navigation within a program fragment, and the ability to actively personalize content via various end-user content enrichments (such as line art, referrals and hyperlink insertions). In this paper, we present the implementation of a range of "couch-top" control and editing devices, including personal devices such as personal digital assistants and ad-hoc interactive devices. This paper also presents an architecture that decouples user actions into activators and handlers. We provide an overview of the interaction architecture and report on a series of deployment experiments on a wide range of consumer electronics devices. Dick C. A. Bulterman, Pablo César, Jack Jansen 0001 |
ACM Multimedia | 2 |
| 2006 | Interactive digital television and multimedia systemsabstractInteractive digital television is an emerging field with a high impact in our societies: it offers interactive services to the masses. This tutorial aims to establish a common framework by summarizing the most significant results in this multidisciplinary field. The review includes topics such as content distribution, system software of the receivers, and user interaction. In addition, we will discuss current commercial events such as the next generation of optical discs (e.g., blue-ray), BBC peer-to-peer service, and mobile television. Based on this discussion, we will formulate an agenda for further research. The agenda includes, for example, end-user enrichment of television content and social television. This half-day tutorial will provide the attendee a solid understanding of the technologies currently in use and an introduction of the open questions in the field. Pablo César, Konstantinos Chorianopoulos |
ACM Multimedia | 1 |
| 2006 | Open graphical framework for interactive TV
Pablo César, Juha P. Vierinen, Petri Vuorimaa |
Multim. Tools Appl. | 1 |
| 2006 | A graphics architecture for high-end interactive television terminalsabstractThis article presents a graphics software architecture for next-generation digital television receivers. We propose that such receivers should include a standardised Java-based procedural environment capable of rendering 2D/3D graphics and video, and a declarative environment supporting W3C recommendations such as SMIL and XForms. We also introduce a graphics architecture model that meets such requirements. As a proof-of-concept, a prototype implementation of the model is presented. This implementation enhances television content by allowing the user to play 3D graphics games, to run Java applications, and to browse XML-based documents while meeting current hardware restrictions. Pablo César, Petri Vuorimaa, Juha P. Vierinen |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2001 | System Software For Digital Television ApplicationsabstractInteractive Television is fast becoming a necessity as it converges the popular web browsing and the standard television systems better. This paper discusses the underlying system - Operating system and Java Runtime Environment - for the Digital TV. A review of the needed system capabilities for Digital TV, a probable solution of the underlying system, and future improvisation of the system are dealt herewith. Ganesh Sivaraman, Pablo César, Petri Vuorimaa |
ICME | 2 |