Jack Jansen 0001

dblp:57/1786 · also A. J. Jansen 0001 · DBLP profile ↗
← Back
49ranked-venue papers
12as first author
15since 2021 · last 2026
0000-0002-7006-2560ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 5 first-author · 14 since 2021Databases, data management, data science and information retrieval · 10 · 6 first-authorHuman-computer interaction and ubiquitous computing · 10 · 6 since 2021Computer networks · 8 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 WebRTC-Based Volumetric Video Conferencing: SFU Architecture Evaluation and Benchmarking
abstract
Immersive technologies promise to revolutionize communication through enhanced sense of presence and interactivity. To enable interaction, reliable low-latency transport mechanisms are needed to handle the large volumes of data created by complex 3D objects. In this paper, we propose an open-source, codec-independent, selective forwarding unit (SFU) for real-time volumetric video streaming using WebRTC. For evaluation purposes, we provide a reference client implementation by extending VR2Gather, a TCP-based system for immersive communication. We conduct extensive evaluations using both new and existing datasets to compare the performance of WebRTC against TCP-based protocols in an emulated testbed environment. The evaluations demonstrate that WebRTC outperforms other protocols in high-latency scenarios and adapts video quality to user movement 13% and 36% faster than its TCP-based counterparts in networks with 5 ms and 10 ms of network latency, respectively.
Matthias De Fré, Jeroen van der Hooft, Jack Jansen 0001, Silvia Rossi 0001, Thomas Röggla, Tim Wauters, Filip De Turck, Irene Viola 0001, Pablo César
NOSSDAV3
2026 Social XR for Pre-Production Meetings: An In-the-Wild Study of Early Stage Communication Between XR Producers and Clients
Sueyoon Lee, Irene Viola 0001, Jack Jansen 0001, Ashutosh Singla, Karolina Wylezek, Thomas Röggla, Pablo César
IMX3
2026 The Influence of Context on Learning in a Social VR Historical Fashion Exhibition
Karolina Wylezek, Irene Viola 0001, Silvia Rossi 0001, Jack Jansen 0001, Thomas Röggla, Pablo César
IMX4
2025 QoE Evaluation of Remote Physiotherapy in Volumetric Video and Video-Based Real-Time Communication
abstract
In recent years, video conferencing platforms have become powerful tools for remote communication. There has also been an increase in the use of VR systems for communication. However, very few of these systems utilize photorealistic human representation. This paper investigates the strengths, challenges, and limitations of a novel 3D communication prototype (VR2Gather) and a well-established video conferencing system (Zoom). Specifically, we explore whether the 3D communication prototype can achieve comparable performance levels in a remote physiotherapy use case. By assessing various aspects, such as audio-visual quality, presence, and interaction, we aim to determine if the current prototype is comparable with commercial systems in some dimensions while exceeding expectations in others. Our results indicated that VR2Gather has the potential for a better sense of connection and higher concentration. However, challenges like improving 3D rendering quality and communication ease still need to be overcome to make it suitable for physiotherapy.
Ashutosh Singla, Irene Viola 0001, Jack Jansen 0001, Pablo César
ICME3
2025 UVG-CWI-DQPC: Dual-Quality Point Cloud Dataset for Volumetric Video Applications
abstract
Volumetric video is a key enabler of immersive extended reality (XR) experiences and is often represented using point clouds for their structural simplicity. However, capturing volumetric content through multi-view acquisition and depth sensing poses many challenges, such as occlusions and depth mismatches. To foster research in this field, we introduce a unique dual-quality point cloud dataset, named UVG-CWI-DQPC, which is designed to support the development of point cloud enhancement, compression, and quality assessment. Our dataset includes 12 dynamic sequences captured simultaneously by: 1) a high-end capture system producing high-fidelity point clouds with extensive processing; and 2) a consumer-grade capture system relying on affordable RGB-D cameras, lightweight processing, and open-source tools. For each sequence, our dataset provides ground-truth point clouds from the high-end capture system and raw RGB-D footage from the consumer-grade capture system, along with calibration data and tools for point cloud generation. This dual-quality setup enables direct comparison and benchmarking of algorithms for densification, occlusion removal, registration, and quality enhancement. Our dataset is publicly available under a permissive license to support reproducible research and standardization work in Moving Picture Experts Group (MPEG) and 3rd Generation Partnership Project (3GPP).
Guillaume Gautier, Xuemei Zhou, Jack Jansen 0001, Louis Fréneau, Marko Viitanen, Uyen Phan, Jani Käpylä, Irene Viola 0001, Alexandre Mercat, Pablo César, Jarno Vanne
ACM Multimedia4
2025 From Individual QoE to Shared Mental Models: A Novel Evaluation Paradigm for Collaborative XR
abstract
Extended Reality (XR) systems are rapidly shifting from isolated, single-user applications towards collaborative and social multi-user experiences. To evaluate the quality and effectiveness of such interactions, it is therefore required to move beyond traditional individual metrics such as Quality-of-Experience (QoE) or Sense of Presence (SoP). Instead, group-level dynamics such as effective communication, coordination etc. need to be encompassed to assess the shared understanding of goals and procedures. In psychology, this is referred to as a Shared Mental Model (SMM). The strength and congruence of such an SMM are known to be key for effective team collaboration and performance. In an immersive XR setting, though, novel Influence Factors (IFs) emerge that are not considered in a setting of physical co-location. Evaluations on the impact of these novel factors on SMM formation in XR, however, are close to non-existent. Therefore, this work proposes SMMs as a novel evaluation tool for collaborative and social XR experiences. To better understand how to explore this construct, we ran a prototypical experiment based on ITU recommendations in which the influence of asymmetric end-to-end latency is evaluated through a collaborative, two-user block building task. The results show how also in an XR context strong SMM formation can take place even when collaborators have fundamentally different responsibilities and behavior. Moreover, the study confirms previous findings by showing in an XR context that a teams’ SMM strength is positively associated with its performance.
Sam Van Damme, Jack Jansen 0001, Silvia Rossi 0001, Pablo César
QoMEX2
2025 Subjective and Objective Quality Assessment for Dynamic Point Cloud with Visual Attention in 6 DoF
abstract
Perceptual quality assessment of Dynamic Point Cloud (DPC) contents plays an important role in various Virtual Reality (VR) applications that involve human beings as the end user. Understanding and modeling perceptual quality assessment is greatly enriched by insights from visual attention. However, incorporating aspects of visual attention in DPC quality models is largely unexplored, as ground-truth visual attention data are scarcely available. Besides, testing methods and procedures for collecting visual attention data are still to be agreed on. This article presents a dataset containing subjective opinion scores and visual attention maps of DPCs, collected in a VR environment using eye-tracking technology. Both the quality score and eye-tracking data were collected during a subjective quality assessment experiment, in which subjects were instructed to watch and rate DPCs at various degradation levels under 6 Degrees of Freedom (DoF) inspection, using a head-mounted display. Qualitative interview analysis was also conducted after the experiment. The dataset consists of 50 DPCs, including 5 reference DPCs, with each reference encoded at 3 distortion levels using 3 different codecs (namely G-PCC, V-PCC, CWI-PCL), amounting to a total of 9 degraded version per reference. Additionally, it incorporates 1,000 gaze trials from 40 participants, yielding a total of 15,000 visual attention maps across all the DPCs. We additionally benchmark objective quality metrics originally designed for static point clouds, evaluating their performance in our dataset using two temporal pooling strategies. Furthermore, we employ the visual attention data that are retrieved during our experiment to evaluate whether the performance of widely used objective quality metrics is improved by considering subjective measurements of visual attention. This dataset establishes a link between quality assessment and visual attention within the context of DPC. Moreover, thematic analysis of the interviews helps uncover user behavior and factors impacting perceptual quality for DPC in 6 DoF. This work deepens our understanding of DPC quality assessment and visual attention, driving progress in the realm of VR experiences and perception.
Xuemei Zhou, Irene Viola 0001, Evangelos Alexiou, Jack Jansen 0001, Pablo César
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Open-Sourcing VR2Gather: A Collaborative Social VR System for Adaptive Multi-Party Real Time Communication
abstract
Social Virtual Reality is envisioned to transform how individu- als communicate remotely, offering a sense of immersion and co- presence within a virtual space. Current platforms enabling remote social interactions rely on synthetic user representations. We ad- dress this limitation by enabling realistic human representation through volumetric content capture, encoding and transmission. Specifically, we present an extended version of VR2Gather, now a fully open source Unity package, available at https://github.com/ cwi-dis/VR2Gather-acmmm-oss. Our platform is a customisable system to transmit volumetric content in a multi-party real-time environment, easy to integrate into existing applications.
Jack Jansen 0001, Thomas Röggla, Silvia Rossi 0001, Irene Viola 0001, Pablo César
ACM Multimedia1
2024 Delay Threshold for Social Interaction in Volumetric eXtended Reality Communication
abstract
Immersive technologies like eXtended Reality (XR) are the next step in videoconferencing. In this context, understanding the effect of delay on communication is crucial. This article presents the first study on the impact of delay on collaborative tasks using a realistic Social XR system. Specifically, we design an experiment and evaluate the impact of end-to-end delays of 300, 600, 900, 1,200, and 1,500 ms on the execution of a standardized task involving the collaboration of two remote users that meet in a virtual space and construct block-based shapes. To measure the impact of the delay in this communication scenario, objective and subjective data were collected. As objective data, we measured the time required to execute the tasks and computed conversational characteristics by analyzing the recorded audio signals. As subjective data, a questionnaire was prepared and completed by every user to evaluate different factors such as overall quality, perception of delay, annoyance using the system, level of presence, cybersickness, and other subjective factors associated with social interaction. The results show a clear influence of the delay on the perceived quality and a significant negative effect as the delay increases. Specifically, the results indicate that the acceptable threshold for end-to-end delay should not exceed 900 ms. This article additionally provides guidelines for developing standardized XR tasks for assessing interaction in Social XR environments.
Carlos Cortés 0001, Irene Viola 0001, Jesús Gutiérrez 0001, Jack Jansen 0001, Shishir Subramanyam, Evangelos Alexiou, Pablo Pérez 0001, Narciso García, Pablo César
ACM Trans. Multim. Comput. Commun. Appl.4
2023 QAVA-DPC: Eye-Tracking Based Quality Assessment and Visual Attention Dataset for Dynamic Point Cloud in 6 DoF
abstract
Perceptual quality assessment of Dynamic Point Cloud (DPC) contents plays an important role in various Virtual Reality (VR) applications that involve human beings as the end user, understanding and modeling perceptual quality assessment is greatly enriched by insights from visual attention. However, incorporating aspects of visual attention in DPC quality models is largely unexplored, as ground-truth visual attention data is scarcely available. This paper presents a dataset containing subjective opinion scores and visual attention maps of DPCs, collected in a VR environment using eye-tracking technology. The data was collected during a subjective quality assessment experiment, in which subjects were instructed to watch and rate DPCs at various degradation levels under 6 degrees-of-freedom inspection, using a head-mounted display. The dataset comprises 5 reference DPC contents, with each reference encoded at 3 distortion levels using 3 different codecs, amounting to a total of 9 degraded DPC contents. Moreover, it includes 1,000 gaze trials from 40 participants, resulting in 15,000 visual attention maps in total. The curated dataset can serve as authentic benchmark data for assessing the performance of objective DPC quality metrics. Additionally, it establishes a link between quality assessment and visual attention within the context of DPC. This work deepens our understanding of DPC quality and visual attention, driving progress in the realm of VR experiences and perception.
Xuemei Zhou, Irene Viola 0001, Evangelos Alexiou, Jack Jansen 0001, Pablo César
ISMAR4
2022 Mediascape XR: A Cultural Heritage Experience in Social VR
abstract
Social virtual reality (VR) allows multiple remote users to interact in a shared space, unveiling new possibilities for communication in immersive environments. Mediascape XR presents a social VR experience that teleports 3D representations of remote users, using volumetric video, to a virtual museum. It enables visitors to interact with cultural heritage artifacts while allowing social interactions in real time between them. The application is designed following a human-centered approach, enabling an interactive, educating, and entertaining experience.
Ignacio Reimat, Yanni Mei, Evangelos Alexiou, Jack Jansen 0001, Jie Li 0064, Shishir Subramanyam, Irene Viola 0001, Johan Oomen, Pablo César
ACM Multimedia4
2022 Evaluating the Impact of Tiled User-Adaptive Real-Time Point Cloud Streaming on VR Remote Communication
abstract
Remote communication has rapidly become a part of everyday life in both professional and personal contexts. However, popular video conferencing applications present limitations in terms of quality of communication, immersion and social meaning. VR remote communication applications offer a greater sense of co-presence and mutual sensing of emotions between remote users. Previous research on these applications has shown that realistic point cloud user reconstructions offer better immersion and communication as compared to synthetic user avatars. However, photorealistic point clouds require a large volume of data per frame and are challenging to transmit over bandwidth-limited networks. Recent research has demonstrated significant improvements to perceived quality by optimizing the usage of bandwidth based on the position and orientation of the user's viewport with user-adaptive streaming. In this work, we developed a real-time VR communication application with an adaptation engine that features tiled user-adaptive streaming based on user behaviour. The application also supports traditional network adaptive streaming. The contribution of this work is to evaluate the impact of tiled user-adaptive streaming on quality of communication, visual quality, system performance and task completion in a functional live VR remote communication system. We performed a subjective evaluation with 33 users to compare the different streaming conditions with a neck exercise training task. As a baseline, we use uncompressed streaming requiring approximately 300 megabits per second and our solution achieves similar visual quality with tiled adaptive streaming at 14 megabits per second. We also demonstrate statistically significant gains in the quality of interaction and improvements to system performance and CPU consumption with tiled adaptive streaming as compared to the more traditional network adaptive streaming.
Shishir Subramanyam, Irene Viola 0001, Jack Jansen 0001, Evangelos Alexiou, Alan Hanjalic, Pablo César
ACM Multimedia3
2022 Subjective QoE Evaluation of User-Centered Adaptive Streaming of Dynamic Point Clouds
abstract
Technological advances in head-mounted displays and novel real-time 3D acquisition and reconstruction solutions have fostered the development of 6 Degrees of Freedom (6DoF) teleimmersive systems for social VR applications. Point clouds have emerged as a popular format for such applications, owing to their simplicity and versatility; yet, dense point cloud contents are too large to deliver directly over bandwidth-limited networks. In this context, user-adaptive delivery mechanisms are a promising solution to exploit the increased range of motion offered by 6DoF VR applications to yield gains in perceived quality of 3D point cloud user representations, while reducing their bandwidth requirements. In this paper, we perform a user study in VR to quantify the gains adaptive tile selection strategies can bring with respect to non-adaptive solutions. In particular, we define an auxiliary utility function, we employ established methods from the literature and newly-proposed schemes for distributing the bit budget across the tiles, and we evaluate them together with non-adaptive streaming baselines through subjective QoE assessment. Results confirm that considerable gains can be obtained with user-adaptive streaming, achieving bit rate gains of up to 65% with respect to a non-adaptive approach to deliver comparable quality. Our analysis provides useful insights for the design and development of social VR applications.
Shishir Subramanyam, Irene Viola 0001, Jack Jansen 0001, Evangelos Alexiou, Alan Hanjalic, Pablo César
QoMEX3
2021 Evaluating the user Experience of a Photorealistic Social VR Movie
abstract
We all enjoy watching movies together. However, this is not always possible if we live apart. While we can remotely share our screens, the experience differs from being together. We present a social Virtual Reality (VR) system that captures, reconstructs, and transmits multiple users’ volumetric representations into a commercially produced 3D virtual movie, so they have the feeling of “being there” together. We conducted a 48-user experiment where we invited users to experience the virtual movie either using a Head Mounted Display (HMD) or using a 2D screen with a game controller. In addition, we invited 14 VR experts to experience both the HMD and the screen version of the movie and discussed their experiences in two focus groups. Our results showed that both end-users and VR experts found that the way they navigated and interacted inside a 3D virtual movie was novel. They also found that the photorealistic volumetric representations enhanced feelings of co-presence. Our study lays the groundwork for future interactive and immersive VR movie co-watching experiences.
Jie Li 0064, Shishir Subramanyam, Jack Jansen 0001, Yanni Mei, Ignacio Reimat, Kinga Lawicka, Pablo César
ISMAR3
2021 CWIPC-SXR: Point Cloud dynamic human dataset for Social XR
abstract
Real-time, immersive telecommunication systems are quickly becoming a reality, thanks to the advances in acquisition, transmission, and rendering technologies. Point clouds in particular serve as a promising representation in these type of systems, offering photorealistic rendering capabilities with low complexity. Further development of transmission, coding, and quality evaluation algorithms, though, is currently hindered by the lack of publicly available datasets that represent realistic scenarios of remote communication between people in real-time. In this paper, we release a dynamic point cloud dataset that depicts humans interacting in social XR settings. Using commodity hardware, we capture a total of 45 unique sequences, according to several use cases for social XR. As part of our release, we provide annotated raw material, resulting point cloud sequences, and an auxiliary software toolbox to acquire, process, encode, and visualize data, suitable for real-time applications. The dataset can be accessed via the following link: https://www.dis.cwi.nl/cwipc-sxr-dataset/.
Ignacio Reimat, Evangelos Alexiou, Jack Jansen 0001, Irene Viola 0001, Shishir Subramanyam, Pablo César
MMSys3
2020 ThermalWear: Exploring Wearable On-chest Thermal Displays to Augment Voice Messages with Affect
abstract
Voice is a rich modality for conveying emotions, however emotional prosody production can be situationally or medically impaired. Since thermal displays have been shown to evoke emotions, we explore how thermal stimulation can augment perception of neutrally-spoken voice messages with affect. We designed ThermalWear, a wearable on-chest thermal display, then tested in a controlled study (N=12) the effects of fabric, thermal intensity, and direction of change. Thereafter, we synthesized 12 neutrally-spoken voice messages, validated (N=7) them, then tested (N=12) if thermal stimuli can augment their perception with affect. We found warm and cool stimuli (a) can be perceived on the chest, and quickly without fabric (4.7-5s) (b) do not incur discomfort (c) generally increase arousal of voice messages and (d) increase / decrease message valence, respectively. We discuss how thermal displays can augment voice perception, which can enhance voice assistants and support individuals with emotional prosody impairments.
Abdallah El Ali, Swamy Ananthanarayan, Thomas Röggla, Jack Jansen 0001, Jessica Hartcher-O'Brien, Kaspar M. B. Jansen, Pablo César
CHI5
2020 A pipeline for multiparty volumetric video conferencing: transmission of point clouds over low latency DASH
abstract
The advent of affordable 3D capture and display hardware is making volumetric videoconferencing feasible. This technology increases the immersion of the participants, breaking the flat restriction of 2D screens, by allowing them to collaborate and interact in shared virtual reality spaces. In this paper we introduce the design and development of an architecture intended for volumetric videoconferencing that provides a highly realistic 3D representation of the participants, based on pointclouds. A pointcloud representation is suitable for real-time applications like video conferencing, due to its low-complexity and because it does not need a time consuming reconstruction process. As transport protocol we selected low latency DASH, due to its popularity and client-based adaptation mechanisms for tiling. This paper presents the architectural design, details the implementation, and provides some referential results. The demo will showcase the system in action, enabling volumetric videoconferencing using pointclouds.
Jack Jansen 0001, Shishir Subramanyam, Romain Bouqueau, Gianluca Cernigliaro, Marc Martos Cabré, Pablo César
MMSys1
2018 Workflow Support for Live Object-Based Broadcasting
abstract
This paper examines the document aspects of object-based broadcasting. Object-based broadcasting augments traditional video and audio broadcast content with additional (temporally-constrained) media objects. The content of these objects -- as well as their temporal validity -- are determined by the broadcast source, but the actual rendering and placement of these objects can be customized to the needs/constraints of the content viewer(s). The use of object-based broadcasting enables a more tailored end-user experience than the one-size-fits-all of traditional broadcasts: the viewer may be able to selectively turn off overlay graphics (such as statistics) during a sports game, or selectively render them on a secondary device. Object-based broadcasting also holds the potential for supporting presentation adaptivity for accessibility or for device heterogeneity.
Jack Jansen 0001, Pablo César, Dick C. A. Bulterman
DocEng1
2017 Tangible Air: An Interactive Installation for Visualising Audience Engagement
abstract
This article presents an end-to-end system for capturing physiological sensor data and visualising it on a real-time graphic dashboard and as part of an art installation. More specifically, it describes an event where the level of engagement of the audience was measured by means of Galvanic Skin Response (GSR) sensors and of the presenter through a sweater fitted with GSR, ECG and acceleration sensors. The gathered data was presented in real-time through a visualisation projected onto a screen and a physical electro-mechanical installation, which would change the height of helium-filled balloons depending on the atmosphere in the auditorium. Thereby trying to create a tangible way of making the invisible visible.
Thomas Röggla, Chen Wang 0034, Lilia Perez Romero, Jack Jansen 0001, Pablo César
Creativity & Cognition4
2017 CWI-ADE2016 Dataset: Sensing nightclubs through 40 million BLE packets
abstract
The CWI-ADE2016 Dataset is a collection of more than 40 million Bluetooth Low Energy (BLE) packets and of 14 million accelerometer and temperature samples generated by wristbands that people wore in a nightclub. The data was gathered during Amsterdam Dance Event 2016 in an exclusive club experience curated around human senses, which leveraged technology as a bridge between the club and the guests. Each guest was handed a custom-made wristband with a BLE-enabled device that broadcast movement, temperature and other sensor readings. A network of Raspberry Pi receivers deployed for the occasion captured broadcast packets from wristbands and any other BLE device in the environment. This data provides a full picture of the performance of the real life deployment of a sensing infrastructure and gives insights to designing sensing platforms, understanding networks and crowds behaviour or studying opportunistic sensing. This paper describes an analysis of this dataset and some examples of usage.
Sergio Cabrero, Jack Jansen 0001, Thomas Röggla, John Alexis Guerra Gómez, David A. Shamma, Pablo César
MMSys2
2017 Co-present and remote audience experiences: intensity and cohesion
Erik Geelhoed, Kuldip Singh-Barmi, Ian Biscoe, Pablo César, Jack Jansen 0001, Chen Wang 0034, Rene Kaiser
Multim. Tools Appl.5
2016 Quantifying audience experience in the wild: Heuristics for developing and deploying a biosensor infrastructure in theaters
abstract
Measuring the experience of audience of arts events is essential in the “experience economy” of this day and age, but it is a difficult task. The value of such information goes beyond evaluating the impact of the arts, as it can provide insights and feedback to enhance the work of artists and the experiences of other audience members. Through in-depth understanding of the needs of the providers and consumers of the arts, we progressively developed a biosensor infrastructure that was deployed in theaters. Over the years, we identified the challenges and issues related to developing and deploying a biosensor infrastructure in theaters. These collective experiences and identified issues were categorized into three main areas: processes, data, and system. A total of seven heuristics are developed across the three main areas. Processes place the stakeholders and audiences at the core of the research; data provides guidelines for data validity, collecting a variety of data, and supporting real-time data gathering; and systems covers the concurrency, scalability, deployment and feedback of the infrastructure. We believe that this set of heuristics forms the foundation for an adequate infrastructure to measure audience experience in the wild and it is a valuable source of guideline for future work.
Chen Wang 0034, Jacqueline Wong, Xintong Zhu, Thomas Röggla, Jack Jansen 0001, Pablo César
QoMEX5
2015 Multimedia Document Structure for Distributed Theatre
abstract
This paper explores the suitability of structured (and declarative) multimedia document formats for supporting a novel type of performing arts: distributed theatre. In distributed theatre, the actors are split between two (or more) locations, but together deliver a single performance mediated by the cameras, the internet, and projection technologies. Based on our efforts to make an actual distributed theatre production happen (the Tempest by Miracle Theatre), this paper reflects on our experience. Our findings are divided into two main areas: workflow and document structure. We conclude that novel types of video-mediated applications, like distributed theatre, require new manners of authoring documents. Moreover, specific extensions to existing document formats are needed in order to accommodate the new requirements imposed by such kind of applications.
Jack Jansen 0001, Michael Frantzis, Pablo César
DocEng1
2015 A Distributed Theatre Experiment with Shakespeare
abstract
This paper reports on an experimental production of The Tempest that was developed in collaboration with Miracle Theatre Company realised as a distributed performance from two separate stages through a dynamically configured telepresence system. The production allowed an exploration of the way a range of technologies, including consumer grade broadband, cameras and projection technologies could affect the development and delivery of live theatre by regional touring company. The architecture of the communication platform used to deliver the performance is introduced as are two novel software tools that are used to describe and control the way the play should be captured and represented.
Doug Williams, Ian Kegel, Marian Florin Ursu, Pablo César, Jack Jansen 0001, Erik Geelhoed, Andras Horti, Michael Frantzis, Bill Scott
ACM Multimedia5
2014 VideoLat: An Extensible Tool for Multimedia Delay Measurements
abstract
When using a videoconferencing system there will always be a delay from sender to receiver. Such delays affect human communication, and therefore knowing the delay is a major factor in judging the expected quality of experience of the conferencing system. Additionally, for implementors, tuning the system to reduce delay requires an ability to effectively and easily gather delay metrics on a potentially wide range of settings. In order to support this process, we make available a system called videoLat. VideoLat provides an innovative approach to understand glass-to-glass video delays and speaker-to-microphone delays. VideoLat can be used as-is to do audio and video roundtrip delays of black box systems, but by making it available as open source we want to enable people to extend and modify it for different scenarios, such as measuring one-way delays or delay of camera switching.
Jack Jansen 0001
ACM Multimedia1
2013 Multimedia document synchronization in a distributed social context
abstract
Watching digital content together and commenting on it is becoming a social habit between friends and family members living apart. It is also becoming an important value-added activity for business video conferencing. In both cases, the video sharing experience can easily be spoiled if synchronization problems arise, since the context of the conversation will not be consistent across locations. In the past, research has treated the distributed synchronization problem as a technical one, mainly focusing on timestamps, frame accuracy, and protocol-dependent control messages. That approach is based on a content agnostic approach which we feel does not adequately address the higher-level constraints of individual conversations.
Jack Jansen 0001, Pablo César, Dick C. A. Bulterman
ACM Symposium on Document Engineering1
2013 User-centric video delay measurements
abstract
The complexities and physical constraints associated with video transmission make the introduction of video playout delays unavoidable. Tuning systems to reduce delay requires an ability to effectively and easily gather delay metrics on a potentially wide range of systems. In order to support this process, we report on a system called videoLat. VideoLat provides an innovative approach to understand glass-to-glass video delays. This paper provides a series of requirements for obtaining representative delay information, it illustrates how such measurements can provide insights into complex (and often closed) video processing systems, and it describes how user-centric testing can be supported in a more realistic manner. We also survey the present state of the art in video delay measurement. The main contribution of this work is that it provides a measuring framework that could serve as the basis for obtaining representative comparative measurements across a wide range of video processing environments.
Jack Jansen 0001, Dick C. A. Bulterman
NOSSDAV1
2012 Just-in-time personalized video presentations
abstract
Using high-quality video cameras on mobile devices, it is relatively easy to capture a significant volume of video content for community events such as local concerts or sporting events. A more difficult problem is selecting and sequencing individual media fragments that meet the personal interests of a viewer of such content. In this paper, we consider an infrastructure that supports the just-in-time delivery of personalized content. Based on user profiles and interests, tailored video mash-ups can be created at view-time and then further tailored to user interests via simple end-user interaction. Unlike other mash-up research, our system focuses on client-side compilation based on personal (rather than aggregate) interests. This paper concentrates on a discussion of language and infrastructure issues required to support just-in-time video composition and delivery. Using a high school concert as an example, we provide a set of requirements for dynamic content delivery. We then provide an architecture and infrastructure that meets these requirements. We conclude with a technical and user analysis of the just-in-time personalized video approach.
Jack Jansen 0001, Pablo César, Rodrigo Laiola Guimarães, Dick C. A. Bulterman
ACM Symposium on Document Engineering1
2012 Enabling 'togetherness' in high-quality domestic video
abstract
Low-cost video conferencing systems have provided an existence proof for the value of video communication in a home setting. At the same time, current systems have a number of fundamental limitations that inhibit more general social interactions among multiple groups of participants. In our work, we describe the development, implementation and evaluation of a domestic video conferencing system that is geared to providing true 'togetherness' among conference participants. We show that such interactions require sophisticated support for high-quality audiovisual presentation, and processing support for person identification and localisation. In this paper, we describe user requirements for effective interpersonal interaction. We then report on a system that implements these requirements. We conclude with a systems and user evaluation of this work. We present results that show that participants in a video conference can be made feel as 'together' as collocated players of a board game.
Ian Kegel, Pablo César, Jack Jansen 0001, Dick C. A. Bulterman, Tim Stevens, Joke Kort, Nikolaus Färber
ACM Multimedia3
2012 A URI-based approach for addressing fragments of media resources on the Web
Erik Mannens, Davy Van Deursen, Raphaël Troncy, Silvia Pfeiffer, Conrad Parker, Yves Lafon, Jack Jansen 0001, Michael Hausenblas, Rik Van de Walle
Multim. Tools Appl.7
2011 Accurate and low-delay seeking within and across mash-ups of highly-compressed videos
abstract
In typical video mash-up systems, a group of source videos are compiled off-line into a single composite object. This improves rendering performance, but limits the possibilities for dynamic composition of personalized content. This paper discusses systems and network issues for enabling client-side dynamic composition of video mash-ups. In particular, this paper describes a novel algorithm to support accurate, low-delay seamless composition of independent clips. We report on an intelligent application-steered scheme that allows system layers to prefetch and discard predicted frames before the rendering moment of indexed content. This approach unifies application-level quality-of-experience specification with system layer quality-of-service processing. To evaluate our scheme, several experiments are conducted and substantial performance improvements are observed in terms of accuracy and low delay.
Jack Jansen 0001, Pablo César, Dick C. A. Bulterman
NOSSDAV2
2011 Enabling Composition-Based Video-Conferencing for the Home
abstract
This paper describes a videoconferencing system that meets performance constraints and functional requirements for use in consumer homes. Our system improves existing home technologies (such as video chat) by providing high-quality audiovisual communication, efficient encoding mechanisms, and low end-to-end delay. Moreover, the system includes a control interface that is capable of dynamically manipulating and compositing audiovisual content streams. This innovative architectural component is required for a domestic setting, where the television acts as the main screen and multiple people gather around it. Apart from the requirements and architecture, this paper analyses the performance of our system. The results validate our architectural decisions and provide a valuable input for further research in domestic videoconferencing.
Jack Jansen 0001, Pablo César, Dick C. A. Bulterman, Tim Stevens, Ian Kegel, J. Issing
IEEE Trans. Multim.1
2010 A model for editing operations on active temporal multimedia documents
abstract
Inclusion of content with temporal behavior in a structured document leads to such a document gaining temporal semantics. If we then allow changes to the document during its presentation, this brings with it a number of fundamental issues that are related to those temporal semantics. In this paper we study modifications of active multimedia documents and the implications of those modifications for temporal consistency. Such modifications are becoming increasingly important as multimedia documents move from being primarily a standalone presentation format to being a building block in a larger application.
Jack Jansen 0001, Pablo César, Dick C. A. Bulterman
ACM Symposium on Document Engineering1
2009 Leveraging user impact: an architecture for secondary screens usage in interactive television
Pablo César, Dick C. A. Bulterman, Jack Jansen 0001
Multim. Syst.3
2009 SMIL State: an architecture and implementation for adaptive time-based web applications
Jack Jansen 0001, Dick C. A. Bulterman
Multim. Tools Appl.1
2009 Fragment, tag, enrich, and send: Enhancing social sharing of video
abstract
The migration of media consumption to personal computers retains distributed social viewing, but only via nonsocial, strictly personal interfaces. This article presents an architecture, and implementation for media sharing that allows for enhanced social interactions among users. Using a mixed-device model, our work allows targeted, personalized enrichment of content. All recipients see common content, while differentiated content is delivered to individuals via their personal secondary screens. We describe the goals, architecture, and implementation of our system in this article. In order to validate our results, we also present results from two user studies involving disjoint sets of test participants.
Pablo César, Dick C. A. Bulterman, Jack Jansen 0001, David Geerts, Hendrik Knoche, William Seager
ACM Trans. Multim. Comput. Commun. Appl.3
2008 Multimedia content transformation: fragmentation, enrichment, and adaptation
abstract
This working session will be an interactive discussion about multimedia content transformation. The basic assumption is that content transformation activities should be provided as non-destructive operations. The final goal of the panel is to gather researchers within the community interested in manipulating multimedia content for providing rich user experiences. The organizers of the panel will moderate and shape the discussion; nevertheless, position papers from the participants are expected.
Pablo César, Dick C. A. Bulterman, Jack Jansen 0001, Maria da Graça Campos Pimentel, Simone D. J. Barbosa
ACM Symposium on Document Engineering3
2008 Enabling adaptive time-based web applications with SMIL state
abstract
In this paper we examine adaptive time-based web applications (or presentations). These are interactive presentations where time dictates the major structure, and that require interactivity and other dynamic adaptation. We investigate the current technologies available to create such presentations and their shortcomings, and suggest a mechanism for addressing these shortcomings. This mechanism, SMIL State, can be used to add user-defined state to declarative time-based languages such as SMIL or SVG animation, thereby enabling the author to create control flows that are difficult to realize within the temporal containment model of the host languages. In addition, SMIL State can be used as a bridging mechanism between languages, enabling easy integration of external components into the web application.
Jack Jansen 0001, Dick C. A. Bulterman
ACM Symposium on Document Engineering1
2008 Enhancing social sharing of videos: fragment, annotate, enrich, and share
abstract
Media consumption is an inherently social activity, serving to communicate ideas and emotions across both small- and large-scale communities. The migration of the media experience to personal computers retains social viewing, but typically only via a non-social, strictly personal interface. This paper presents an architecture and implementation for media content selection, content (re)organization, and content sharing within a user community that is heterogeneous in terms of both participants and devices. In addition, our application allows the user to enrich the content as a differentiated personalization activity targeted to his/her peer-group. We describe the goals, architecture and implementation of our system in this paper. In order to validate our results, we also present results from two user studies involving disjoint sets of test participants.
Pablo César, Dick C. A. Bulterman, David Geerts, Jack Jansen 0001, Hendrik Knoche, William Seager
ACM Multimedia4
2008 A presentation layer mechanism for multimedia playback mobility in service oriented architectures
abstract
This paper presents an approach for media presentation continuity in playback mode. We use the term presentation continuity over session transfer since our solution is at the presentation layer. Previous research on this topic has focused on transferring a particular stream or set of related streams at the sessions layer for live broadcasting or conferencing sessions. We argue that in the realm of service oriented architectures, such as telecom operator networks, this approach does not take full advantage of the particular case of media playback. Our mechanism presents an alternative to the traditional approach, which i) Lowers network control plane overhead, thus reducing chances of presentation consistency loss ii) Lowers network data overhead due to lesser need for transcoding iii) Delegates presentation consistency issues, such as inter-media synchronisation, to the media player iv) Dynamically adapts the presentation to the new target devices without transcoding. Finally, we present experimental results which show that our approach is implementable in acceptable time bounds.
Ishan Vaishnavi, Pablo César, Jack Jansen 0001, Dick C. A. Bulterman
MUM3
2008 A mechanism for presentation-layer media continuity in media playback mode
abstract
This demo presents a new approach for media presentation continuity in playback mode. We use the term presentation continuity over session transfer since our solution is at the presentation layer. Previous research on this topic has focused on transferring a particular stream or set of related streams at the sessions layer. Our approach presents an alternative, recognising the fact that a user is connected to a media presentation, which, may be composed of multiple sessions. The advantages of our approach are i) Lower network control plane overhead, thus reducing chances of semantic presentation loss ii) Lower network data overhead due to lesser need for transcoding iii) delegating presentation semantic issues, such as inter-media synchronisation, to the player iv) dynamically adapt the presentation to the new target devices without transcoding.
Ishan Vaishnavi, Pablo César, Jack Jansen 0001, Dick C. A. Bulterman
NOSSDAV3
2007 An efficient, streamable text format for multimedia captions and subtitles
abstract
In spite of the high profile of media types such as video, audio and images, many multimedia presentations rely extensively on text content. Text can be used for incidental labels, or as subtitles or captions that accompany other media objects. In a multimedia document, text content is not only constrained by the need to support presentation styles and layout, it is also constrained by the temporal context of the presentation. This involves intra-text and extra text timing synchronization with other media objects. This paper describes a new timed-text representation language that is intended to be embedded in a non-text host language. Our format, which we call aText (for the Ambulant Text Format), balances the need for text styling with the requirement for an efficient representation that can be easily parsed and scheduled at runtime. aText, which can also be streamed, is defined as an embeddable text format for use within declarative XML languages. The paper presents a discussion of the requirements for the format, a description of the format and a comparison with other existing and emerging text formats. We also provide examples for aText when embedded within the SMIL and MLIF languages and discuss our implementation experiences of aText with the Ambulant Player.
Dick C. A. Bulterman, Jack Jansen 0001, Pablo César, Samuel Cruz-Lara
ACM Symposium on Document Engineering2
2006 Benefits of structured multimedia documents in IDTV: the end-user enrichment system
abstract
This paper presents a system that exploits the benefits of modelling multimedia presentations as structured documents within the context of interactive digital television systems. Our work permits end-users to easily enrich multimedia content at viewing time (e.g., add images and delete scenes). Because the document is structured, the system can expose to the user the possible enrichment alternatives depending on the current state of the presentation (e.g., current story). Moreover, because the base content is wrapped as a structured document, the enrichments can be modelled as overlying layers that do not alter the original content. Finally, the user can share the enriched content (or parts of it) to specific peers within a P2P network.
Pablo César, Dick C. A. Bulterman, Jack Jansen 0001
ACM Symposium on Document Engineering3
2006 The ambulant annotator: empowering viewer-side enrichment of multimedia content
abstract
This paper presents a set of demos that allow viewer-side enrichment of multimedia content in a home setting. The most relevant features of our system are the following: passive authoring of content in contraposition to the traditional active PC authoring, preservation of the base content, and collaborative authoring (e.g., to share the enriched material with a peer group). These requirements are met by modelling television content as structured multimedia documents using SMIL 2.1.
Pablo César, Dick C. A. Bulterman, Jack Jansen 0001
ACM Symposium on Document Engineering3
2006 An architecture for viewer-side enrichment of TV content
abstract
This paper presents a user interface model and implementation for exploiting next-generation interactive capabilities with the domain of television content. Our work studies capabilities that extend a user's potential impact over the consumption and sharing of television programs. The main capabilities of our environment include personalized viewing and navigation within a program fragment, and the ability to actively personalize content via various end-user content enrichments (such as line art, referrals and hyperlink insertions). In this paper, we present the implementation of a range of "couch-top" control and editing devices, including personal devices such as personal digital assistants and ad-hoc interactive devices. This paper also presents an architecture that decouples user actions into activators and handlers. We provide an overview of the interaction architecture and report on a series of deployment experiments on a wide range of consumer electronics devices.
Dick C. A. Bulterman, Pablo César, Jack Jansen 0001
ACM Multimedia3
2004 Ambulant: a fast, multi-platform open source SMIL player
abstract
This paper provides an overview of the Ambulant Open SMIL player. Unlike other SMIL implementations, the Ambulant Player is a reconfigureable SMIL engine that can be customized for use as an experimental media player core. The Ambulant Player is a reference SMIL engine that can be integrated in a wide variety of media player projects. This paper starts with an overview of our motivations for creating a new SMIL engine, then discusses the architecture of the Ambulant Core (including the scalability and custom integration features of the player). We close with a discussion of our implementation experiences with Ambulant instances for Windows, Mac and Linux versions for desktop and PDA devices.
Dick C. A. Bulterman, Jack Jansen 0001, Kleanthis Kleanthous, Kees Blom, Daniel Benden
ACM Multimedia2
1998 GRiNS: A GRaphical INterface for Creating and Playing SMIL Documents
Dick C. A. Bulterman, Lynda Hardman, Jack Jansen 0001, K. Sjoerd Mullender, Lloyd Rutledge
Comput. Networks3
1993 CMIFed: A Presentation Environment for Portable Hypermedia Documents
abstract
Article CMIFed: a presentation environment for portable hypermedia documents Share on Authors: Guido van Rossum View Profile , Jack Jansen View Profile , K. Sjoerd Mullender View Profile , Dick C. A. Bulterman View Profile Authors Info & Claims MULTIMEDIA '93: Proceedings of the first ACM international conference on MultimediaSeptember 1993 Pages 183–188https://doi.org/10.1145/166266.166287Published:01 September 1993 58citation550DownloadsMetricsTotal Citations58Total Downloads550Last 12 Months19Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Guido van Rossum, Jack Jansen 0001, K. Sjoerd Mullender, Dick C. A. Bulterman
ACM Multimedia2
1992 Replication techniques for speeding up parallel applications on distributed systems
abstract
Abstract Most methods for programming loosely coupled systems are based on message‐passing. Recently, however, methods have emerged based on ‘virtually’ sharing data. These methods simplify distributed programming, but are hard to implement efficiently, as loosely coupled systems do not contain physical shared memory. We introduce a new model,the shared data‐object model, that eases the implementation of parallel applications on loosely coupled systems, but can still be implemented efficiently. In our model, shared data are encapsulated in passive data‐objects, which are variables of user‐defined abstract data types. To speed up access to shared data, data‐objects are replicated. This ability to replicate objects is a significant difference with other object‐based models (e.g. Emerald and Amber). Also, by replicating logical objects rather than physical pages, our model has many advantages over shared virtual memory systems. This paper discusses the design choices involved in replicating objects and their effect on performance. Important issues are: how to maintain consistency among different copies of an object; how to implement changes to objects; which strategy for object replication to use. We have implemented several options to determine which ones are the most efficient.
Henri E. Bal, M. Frans Kaashoek, Andrew S. Tanenbaum, Jack Jansen 0001
Concurr. Pract. Exp.4