Jean-Marie Normand

dblp:86/6910 · DBLP profile ↗
← Back
24ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0003-0557-4356ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 10 since 2021Human-computer interaction and ubiquitous computing · 8 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 ImmersiveDMT: A VR System for Expressive Body Movement and Stress Reduction through Dance Movement Therapy
abstract
In our daily lives, we are exposed to various stressors, which, even in mild forms, can lead to serious health issues when prolonged. Dance Movement Therapy (DMT) uses bodily movement as psychotherapy to improve health and well-being; however, traditional DMT requires therapists, specific setups, and physical props, limiting daily accessibility. To address this, we conducted a formative study with 10 DMT experts and developed ImmersiveDMT, a VR application that provides real-time visual and auditory feedback in response to users’ movements. A user study with 56 participants examined how visual, auditory, and audiovisual feedback affect expressive movement and stress reduction compared to a non-responsive immersive baseline. Results showed that auditory and audiovisual conditions induced significantly more expressive movement compared to baseline. Both visual and auditory feedback achieved significant stress reduction below baseline (V: 17.1%, p=.004; A: 16.6%, p=.042), while combined audiovisual feedback led to less stress reduction. This research presents a novel approach to accessible daily stress management by integrating VR and DMT, demonstrating new possibilities for self-directed mental health support.
Satomi Tokida, Yong-Hao Hu, Yuta Itoh 0001, Jean-Marie Normand, Rebecca Fribourg, Jean-Philippe Rivière, Yuichi Hiroi, Takefumi Hiraki, Yoshio Ishiguro
VR4
2025 Measuring the Impact of Objects' Physicalization, Avatar Appearance, and their Consistency on Pick-and-Place Performance in Augmented Reality
abstract
Augmented Reality (AR) is a growing technology that enables interaction with both virtual and real objects. However, in order to support the future development of efficient and usable AR interactions, there is still a lack of systematic knowledge establishing basic interaction performance across different conditions. Therefore, in this paper, we report a user study measuring the impact of objects' physicalization (object's set composed of (i) virtual, (ii) real, or (iii) a composite mix of real and virtual objects) and hand appearance (hand's appearance displayed as (i) the real hand, (ii) an avatar, or (iii) dynamically adapting to the surrounding objects' physicalization) on the speed performance of a pick-and-place task. Overall, our results reveal that objects' physicalization plays a significant role in interaction performance, with the more real objects in a set the better the performance. Moreover, our results also suggest that pick-and-place interaction performances are mostly unaffected by the hand appearance. Interestingly, we also observed that interactions with real objects were less efficient as the object condition required the user to alternate between interactions with virtual and real objects (object condition (iii)), which provides novel insights into an important - mostly AR-specific - factor to consider for designing future AR interactions. Taken together, our results provide a rich characterization of different factors influencing different phases of a pick-and-place interaction, which could be employed to improve the design of future AR applications.
Antonin Cheymol, Jacob Wallace, Juri Yoneyama, Rebecca Fribourg, Jean-Marie Normand, Ferran Argelaguet
IEEE Trans. Vis. Comput. Graph.5
2024 Avatar-Centered Feedback: Dynamic Avatar Alterations Can Induce Avoidance Behaviors to Virtual Dangers
abstract
Abstract One singularity of VR is its capacity to generate a strong sensation of being present in a dangerous environment, without risking the physical consequences. However, this absence of consequences when interacting with virtual dangers might also limit the induction of realistic responses, such as avoidance behaviors, which are a key factor of various VR applications (e.g., training, journalism, or exposure therapies). To address this limitation, we propose avatar-centered feedback, a novel, device-free approach, consisting of dynamically altering the avatar’s appearance and movements to offer coherent feedback from interactions with the virtual environment, including dangers. To begin with, we present a design space clarifying the range of potential implementations for this approach. Then, we tested this approach in a metallurgy scenario, where participants’ virtual hands would redden and display burns as they got closer to a virtual fire (appearance alteration), and simulate short withdrawal reflexes movements when a spark burst next to them (movement alteration). Our results show that in comparison to a control group, participants receiving avatar-centered feedback demonstrated significantly more avoidance behaviors to the virtual fire. Interestingly, we also found that experiencing avatar-centered feedback of fire significantly increased avoidance behaviors toward a following danger of different nature (a circular saw). These results suggest that avatar-centered feedback can also impact the general perception of the avatar vulnerability to the virtual environment.
Antonin Cheymol, Rebecca Fribourg, Anatole Lécuyer, Jean-Marie Normand, Ferran Argelaguet
ISMAR4
2024 Exploring the Influence of Virtual Avatar Heads in Mixed Reality on Social Presence, Performance and User Experience in Collaborative Tasks
abstract
In Mixed Reality (MR), users' heads are largely (if not completely) occluded by the MR Head-Mounted Display (HMD) they are wearing. As a consequence, one cannot see their facial expressions and other communication cues when interacting locally. In this paper, we investigate how displaying virtual avatars' heads on-top of the (HMD-occluded) heads of participants in a Video See-Through (VST) Mixed Reality local collaborative task could improve their collaboration as well as social presence. We hypothesized that virtual heads would convey more communicative cues (such as eye direction or facial expressions) hidden by the MR HMDs and lead to better collaboration and social presence. To do so, we conducted a between-subject study ($\mathrm{n}=88$) with two independent variables: the type of avatar (CartoonAvatar/RealisticAvatar/NoAvatar) and the level of facial expressions provided (HighExpr/LowExpr). The experiment involved two dyadic communication tasks: (i) the "20-question" game where one participant asks questions to guess a hidden word known by the other participant and (ii) a urban planning problem where participants have to solve a puzzle by collaborating. Each pair of participants performed both tasks using a specific type of avatar and facial animation. Our results indicate that while adding an avatar's head does not necessarily improve social presence, the amount of facial expressions provided through the social interaction does have an impact. Moreover, participants rated their performance higher when observing a realistic avatar but rated the cartoon avatars as less uncanny. Taken together, our results contribute to a better understanding of the role of partial avatars in local MR collaboration and pave the way for further research exploring collaboration in different scenarios, with different avatar types or MR setups.
Théo Combe, Rebecca Fribourg, Lucas Detto, Jean-Marie Normand
IEEE Trans. Vis. Comput. Graph.4
2023 Deep weathering effects
abstract
Weathering phenomena are ubiquitous in urban environments, where it is easy to observe severely degraded old buildings as a result of water penetration. Despite being an important part of any realistic city, this kind of phenomenon has received little attention from the Computer Graphics community compared to stains resulting from biological or flow effects on the building exteriors. In this paper, we present physically-inspired deep weathering effects, where the penetration of humidity (i.e., water particles) and its interaction with a building’s internal structural elements result in large, visible degradation effects. Our implementation is based on a particle-based propagation model for humidity propagation, coupled with a spring-based interaction simulation that allows chemical interactions, like the formation of rust, to deform and destroy a building’s inner structure. To illustrate our methodology, we show a collection of deep degradation effects applied to urban models involving the creation of rust or of ice within walls.
Adrien Verhulst, Jean-Marie Normand, Guillaume Moreau, Gustavo Patow
Comput. Graph.2
2023 Beyond my Real Body: Characterization, Impacts, Applications and Perspectives of "Dissimilar" Avatars in Virtual Reality
abstract
In virtual reality, the avatar - the user's digital representation - is an important element which can drastically influence the immersive experience. In this paper, we especially focus on the use of "dissimilar" avatars i.e., avatars diverging from the real appearance of the user, whether they preserve an anthropomorphic aspect or not. Previous studies reported that dissimilar avatars can positively impact the user experience, in terms for example of interaction, perception or behaviour. However, given the sparsity and multi-disciplinary character of research related to dissimilar avatars, it tends to lack common understanding and methodology, hampering the establishment of novel knowledge on this topic. In this paper, we propose to address these limitations by discussing: (i) a methodology for dissimilar avatars characterization, (ii) their impacts on the user experience, (iii) their different fields of application, and finally, (iv) future research direction on this topic. Taken together, we believe that this paper can support future research related to dissimilar avatars, and help designers of VR applications to leverage dissimilar avatars appropriately.
Antonin Cheymol, Rebecca Fribourg, Anatole Lécuyer, Jean-Marie Normand, Ferran Argelaguet
IEEE Trans. Vis. Comput. Graph.4
2023 Handwriting for Efficient Text Entry in Industrial VR Applications: Influence of Board Orientation and Sensory Feedback on Performance
abstract
Text entry in Virtual Reality (VR) is becoming an increasingly important task as the availability of hardware increases and the range of VR applications widens. This is especially true for VR industrial applications where users need to input data frequently. Large-scale industrial adoption of VR is still hampered by the productivity gap between entering data via a physical keyboard and VR data entry methods. Data entry needs to be efficient, easy-to-use and to learn and not frustrating. In this paper, we present a new data entry method based on handwriting recognition (HWR). Users can input text by simply writing on a virtual surface. We conduct a user study to determine the best writing conditions when it comes to surface orientation and sensory feedback. This feedback consists of visual, haptic, and auditory cues. We find that using a slanted board with sensory feedback is best to maximize writing speeds and minimize physical demand. We also evaluate the performance of our method in terms of text entry speed, error rate, usability and workload. The results show that handwriting in VR has high entry speed, usability with little training compared to other controller-based virtual text entry techniques. The system could be further improved by reducing high error rates through the use of more efficient handwriting recognition tools. In fact, the total error rate is 9.28% in the best condition. After 40 phrases of training, participants reach an average of 14.5 WPM, while a group with high VR familiarity reach 16.16 WPM after the same training. The highest observed textual data entry speed is 21.11 WPM.
Nicolas Fourrier, Guillaume Moreau, Mustapha Benaouicha, Jean-Marie Normand
IEEE Trans. Vis. Comput. Graph.4
2023 A Systematic Review of Navigation Assistance Systems for People With Dementia
abstract
Technological developments provide solutions to alleviate the tremendous impact on the health and autonomy due to the impact of dementia on navigation abilities. We systematically reviewed the literature on devices tested to provide assistance to people with dementia during indoor, outdoor and virtual navigation (PROSPERO ID number: 215585). Medline and Scopus databases were searched from inception. Our aim was to summarize the results from the literature to guide future developments. Twenty-three articles were included in our study. Three types of information were extracted from these studies. First, the types of navigation advice the devices provided were assessed through: (i) the sensorial modality of presentation, e.g., visual and tactile stimuli, (ii) the navigation content, e.g., landmarks, and (iii) the timing of presentation, e.g., systematically at intersections. Second, we analyzed the technology that the devices were based on, e.g., smartphone. Third, the experimental methodology used to assess the devices and the navigation outcome was evaluated. We report and discuss the results from the literature based on these three main characteristics. Finally, based on these considerations, recommendations are drawn, challenges are identified and potential solutions are suggested. Augmented reality-based devices, intelligent tutoring systems and social support should particularly further be explored.
Léa Pillette, Guillaume Moreau, Jean-Marie Normand, Manon Perrier, Anatole Lécuyer, Mélanie Cogné
IEEE Trans. Vis. Comput. Graph.3
2022 Studying the Role of Self and External Touch in the Appropriation of Dysmorphic Hands
abstract
In Virtual Reality, self-touch (ST) stimulation is a promising method of sense of body ownership (SoBO) induction that does not require an external effector. However, its applicability to dysmorphic bodies has not been explored yet and remains uncertain due to the requirement to provide incongruent visuomotor sensations. In this, paper, we studied the effect of ST stimulation on dysmorphic hands via haptic retargeting, as compared to a classical external-touch (ET) stimulation, on the SoBO. Our results indicate that ST can induce similar levels of dysmorphic SoBO than ET stimulation, but that some types of dysmorphism might decrease the ST stimulation accuracy due to the nature of the re-targeting that they induce.
Antonin Cheymol, Rebecca Fribourg, Nami Ogawa, Anatole Lécuyer, Yutaro Hirao, Takuji Narumi, Ferran Argelaguet, Jean-Marie Normand
ISMAR8
2022 "Kapow!": Studying the Design of Visual Feedback for Representing Contacts in Extended Reality
abstract
In absence of haptic feedback, the perception of contact with virtual objects can rapidly become a problem in extended reality (XR) applications. XR developers often rely on visual feedback to inform the user and display contact information. However, as for today, there is no clear path on how to design and assess such visual techniques. In this paper, we propose a design space for the creation of visual feedback techniques meant to represent contact with virtual surfaces in XR. Based on this design space, we conceived a set of various visual techniques, including novel approaches based on onomatopoeia and inspired by cartoons, or visual effects based on physical phenomena. Then, we conducted an online preliminary user study with 60 participants, consisting in assessing 6 visual feedback techniques in terms of user experience. We could notably assess, for the first time, the potential influence of the interaction context by comparing the participants’ answers in two different scenarios: industrial versus entertainment conditions. Taken together, our design space and initial results could inspire XR developers for a wide range of applications in which the augmentation of contact seems prominent, such as for vocational training, industrial assembly/maintenance, surgical simulation, videogames, etc.
Julien Cauquis, Victor Mercado, Géry Casiez, Jean-Marie Normand, Anatole Lécuyer
VRST4
2020 Can Retinal Projection Displays Improve Spatial Perception in Augmented Reality?
abstract
Commonly used Head Mounted Displays (HMDs) in Augmented Reality (AR), namely Optical See-Through (OST) displays, suffer from a main drawback: their focal lenses can only provide a fixed focal distance. Such a limitation is suspected to be one of the main factors for distance misperception in AR. In this paper, we studied the use of an emerging new kind of AR display to tackle such perception issues: Retinal Projection Displays (RPDs). With RPDs, virtual images have no focal distance and the AR content is always in focus. We conducted the first reported experiment evaluating egocentric distance perception of observers using Retinal Projection Displays. We compared the precision and accuracy of the depth estimation between real and virtual targets, displayed by either OST HMDs or RPDs. Interestingly, our results show that RPDs provide depth estimates in AR closer to real ones compared to OST HMDs. Indeed, the use of an OST device was found to lead to an overestimation of the perceived distance by 16%, whereas the distance overestimation bias dropped to 4% with RPDs. Besides, the task was reported with the same level of difficulty and no difference in precision. As such, our results shed the first light on retinal projection displays' benefits in terms of user's perception in Augmented Reality, suggesting that RPD is a promising technology for AR applications in which an accurate distance perception is required.
Etienne Peillard, Yuta Itoh 0001, Guillaume Moreau, Jean-Marie Normand, Anatole Lécuyer, Ferran Argelaguet
ISMAR4
2019 Studying Exocentric Distance Perception in Optical See-Through Augmented Reality
abstract
While perceptual biases have been widely investigated in Virtual Reality (VR), very few studies have considered the challenging environment of Optical See-through Augmented Reality (OST-AR). Moreover, regarding distance perception, existing works mainly focus on the assessment of egocentric distance perception, i.e. distance between the observer and a real or a virtual object. In this paper, we study exocentric distance perception in AR, hereby considered as the distance between two objects, none of them being directly linked to the user. We report a user study (n=29) aiming at estimating distances between two objects lying in a frontoparallel plane at 2.1m from the observer (i.e. in the medium-field perceptual space). Four conditions were tested in our study: real objects on the left and on the right of the participant (called real-real), virtual objects on both sides (virtual-virtual), a real object on the left and a virtual one on the right (real-virtual) and finally a virtual object on the left and a real object on the right (virtual-real). Participants had to reproduce the distance between the objects by spreading two real identical objects presented in front of them. The main findings of this study are the overestimation (20%) of exocentric distances for all tested conditions. Surprisingly, the real-real condition was significantly more overestimated (by about 4%, p=.0166) compared to the virtual-virtual condition, i.e. participants obtained better estimates of the exocentric distance for the virtual-virtual condition. Finally, for the virtual-real/real-virtual conditions, the analysis showed a non-symmetrical behavior, which suggests that the relationship between real and virtual objects with respect to the user might be affected by other external factors. Considered together, these unexpected results illustrate the need for additional experiments to better understand the perceptual phenomena involved in exocentric distance perception with real and virtual objects.
Etienne Peillard, Ferran Argelaguet, Jean-Marie Normand, Anatole Lécuyer, Guillaume Moreau
ISMAR3
2019 Virtual Objects Look Farther on the Sides: The Anisotropy of Distance Perception in Virtual Reality
abstract
The topic of distance perception has been widely investigated in Virtual Reality (VR). However, the vast majority of previous work mainly focused on distance perception of objects placed in front of the observer. Then, what happens when the observer looks on the side? In this paper, we study differences in distance estimation when comparing objects placed in front of the observer with objects placed on his side. Through a series of four experiments (n=85), we assessed participants' distance estimation and ruled out potential biases. In particular, we considered the placement of visual stimuli in the field of view, users' exploration behavior as well as the presence of depth cues. For all experiments a two-alternative forced choice (2AFC) standardized psychophysical protocol was employed, in which the main task was to determine the stimuli that seemed to be the farthest one. In summary, our results showed that the orientation of virtual stimuli with respect to the user introduces a distance perception bias: objects placed on the sides are systematically perceived farther away than objects in front. In addition, we could observe that this bias increases along with the angle, and appears to be independent of both the position of the object in the field of view as well as the quality of the virtual scene. This work sheds a new light on one of the specificities of VR environments regarding the wider subject of visual space theory. Our study paves the way for future experiments evaluating the anisotropy of distance perception in real and virtual environments.
Etienne Peillard, Thomas Thebaud, Jean-Marie Normand, Ferran Argelaguet, Guillaume Moreau, Anatole Lécuyer
VR3
2017 A study on the use of an immersive virtual reality store to investigate consumer perceptions and purchase behavior toward non-standard fruits and vegetables
abstract
In this paper we present an immersive virtual reality user study aimed at investigating how customers perceive and if they would purchase non-standard (i.e. misshaped) fruits and vegetables (FaVs) in supermarkets and hypermarkets. Indeed, food waste is a major issue for the retail sector and a recent trend is to reduce it by selling non-standard goods. An important question for retailers relates to the FaVs' “level of abnormality” that consumers would agree to buy. However, this question cannot be tackled using “classical” marketing techniques that perform user studies within real shops since fresh produce such as FaVs tend to rot rapidly preventing studies to be repeatable or to be run for a long time. In order to overcome those limitations, we created a virtual grocery store with a fresh FaVs section where 142 participants were immersed using an Oculus Rift DK2 HMD. Participants were presented either “normal”, “slightly misshaped”, “misshaped” or “severely misshaped” FaVs. Results show that participants tend to purchase a similar number of FaVs whatever their deformity. Nevertheless participants' perceptions of the quality of the FaV depend on the level of abnormality.
Adrien Verhulst, Jean-Marie Normand, Cindy Lombard, Guillaume Moreau
VR2
2016 Real-Time Surface of Revolution Reconstruction on Dense SLAM
abstract
We present a fast and accurate method for reconstructing surfaces of revolution (SoR) on 3D data and its application to structural modeling of a cluttered scene in real-time. To estimate a SoR axis, we derive an approximately linear cost function for fast convergence. Also, we design a framework for reconstructing SoR on dense SLAM. In the experiment results, we show our method is accurate, robust to noise and runs in real-time.
Hideaki Uchiyama, Jean-Marie Normand, Guillaume Moreau, Hajime Nagahara, Rin-Ichiro Taniguchi
3DV3
2016 Practical and Precise Projector-Camera Calibration
abstract
Projectors are important display devices for large scale augmented reality applications. However, precisely calibrating projectors with large focus distances implies a trade-off between practicality and accuracy. People either need a huge calibration board or a precise 3D model [12]. In this paper, we present a practical projectorcamera calibration method to solve this problem. The user only needs a small calibration board to calibrate the system regardless of the focus distance of the projector. Results show that the rootmean-squared re-projection error (RMSE) for a 450cm projection distance is only about 4mm, even though it is calibrated using a small B4 (250×353mm) calibration board.
Jean-Marie Normand, Guillaume Moreau
ISMAR2
2015 Local Geometric Consensus: A General Purpose Point Pattern-Based Tracking Algorithm
abstract
We present a method which can quickly and robustly match 2D and 3D point patterns based on their sole spatial distribution, but it can also handle other cues if available. This method can be easily adapted to many transformations such as similarity transformations in 2D/3D, and affine and perspective transformations in 2D. It is based on local geometric consensus among several local matchings and a refinement scheme. We provide two implementations of this general scheme, one for the 2D homography case (which can be used for marker or image tracking) and one for the 3D similarity case. We demonstrate the robustness and speed performance of our proposal on both synthetic and real images and show that our method can be used to augment any (textured/textureless) planar objects but also 3D objects.
Jean-Marie Normand, Guillaume Moreau
IEEE Trans. Vis. Comput. Graph.2
2014 Robust random dot markers: towards augmented unprepared maps with pure geographic features
abstract
Augmented maps have many important applications. However, no mature registration method exists to associate unprepared maps with a Geographical Information System (GIS) database which would be used to superimpose simulation results or route display on a paper map. In this paper, we propose a method called Robust Random Dot Markers (RRDM) that can robustly track coplanar random dot patterns, which can be used to address this problem. RRDM is based on the same idea of the Random Dot Markers (RDM) proposed by [Uchiyama and Saito 2011] and it can serve as fiducial markers as well as texture independent "natural markers". We conduct a series of experiments and show that RRDM is more robust than RDM in terms of jitter, perspective distortion, under and over detection of dots in the pattern. As an example of "natural marker", we show that RRDM can successfully register unprepared printed maps only with pure geographic features, i.e. road intersections coordinates, which we retrieve from a GIS. Our method does not suffer from the drawbacks of traditional "feature-point" based registration methods which mainly based on textures, since textures may change according to different maps.
Jean-Marie Normand, Guillaume Moreau
VRST2
2013 Real time whole body motion mapping for avatars and robots
abstract
We describe a system that allows for controlling different robots and avatars from a real time motion stream. The underlying problem is that motion data from tracking systems is usually represented differently to the motion data required to drive an avatar or a robot: there may be different joints, motion may be represented by absolute joint positions and rotations or by a root position, bone lengths and relative rotations in the skeletal hierarchy. Our system resolves these issues by remapping in real time the tracked motion so that the avatar or robot performs motions that are visually close to those of the tracked person. The mapping can also be reconfigured interactively at run-time. We demonstrate the effectiveness of our system by case studies in which a tracked person is embodied as an avatar in immersive virtual reality or as a robot in a remote location. We show this with a variety of tracking systems, humanoid avatars and robots.
Bernhard Spanlang, Xavi Navarro, Jean-Marie Normand, Sameer Kishore, Rodrigo Pizarro, Mel Slater
VRST3
2010 A first person avatar system with haptic feedback
abstract
We describe a system that shows how to substitute a person's body in virtual reality by a virtual body (or avatar). The avatar is seen from a first person perspective, moves as the person moves and the system generates touch on the real person's body when the avatar is touched. Such replacement of the person's real body by a virtual body requires a wide field-of-view head-mounted display, real-time whole body tracking, and tactile feedback. We show how to achieve this with a variety of off-the-shelf hardware and software, and also custom systems for real-time avatar rendering and collision detection. We present an overview of the system and detail on some of its components. We provide examples of how such a system is being used in some of our current experimental studies of embodiment.
Bernhard Spanlang, Jean-Marie Normand, Elias Giannopoulos, Mel Slater
VRST2
2008 A Branch and Bound Algorithm for Numerical MAX-CSP
Jean-Marie Normand, Alexandre Goldsztejn, Marc Christie, Frédéric Benhamou
CP1
2008 A Tabu Search Method for Interval Constraints
Charlotte Truchet, Marc Christie, Jean-Marie Normand
CPAIOR3
2008 Camera Control in Computer Graphics
abstract
Abstract Recent progress in modelling, animation and rendering means that rich, high fidelity virtual worlds are found in many interactive graphics applications. However, the viewer's experience of a 3D world is dependent on the nature of the virtual cinematography, in particular, the camera position, orientation and motion in relation to the elements of the scene and the action. Camera control encompasses viewpoint computation, motion planning and editing. We present a range of computer graphics applications and draw on insights from cinematographic practice in identifying their different requirements with regard to camera control. The nature of the camera control problem varies depending on these requirements, which range from augmented manual control (semi‐automatic) in interactive applications, to fully automated approaches. We review the full range of solution techniques from constraint‐based to optimization‐based approaches, and conclude with an examination of occlusion management and expressiveness in the context of declarative approaches to camera control.
Marc Christie, Patrick Olivier, Jean-Marie Normand
Comput. Graph. Forum3
2005 A Semantic Space Partitioning Approach to Virtual Camera Composition
abstract
Positioning a virtual camera in a 3D virtual environment is generally a non-intuitive task when using a 2D input device such as a mouse. The user has a mental representation of the result in terms of what he wants to see on the screen, and has to apply a mental inversion process to determine the location, orientation and field of view parameters of the camera. This task is generally achieved through a tedious and time-consuming process requiring a succession of "place the camera" and "check the result" operations. Current 3D modelers surprisingly lack integration of tools to assist the user in this task, despite the fact that cinema, in more than a hundred years, has provided a rich grammar that allows a director to unambiguously describe shots. Modelers are based on complex mathematical notions (spline curves, velocity graphs) more or less hidden by high-level manipulators. Manipulators allow positioning and animating the camera, but lack correlation with well established cinematographic notions relative to camera composition (object framing, distance shot specification, relative viewing angles, occlusions). In this paper we propose a new approach to virtual camera composition that fully relies on this grammar, relieves the user from low-level parameter manipulation and offers him classes of possible solutions w.r.t. current cinematographic notions. In related literature, numerous approaches have utilized cinematographic properties such as subject size and location within the frame to assist users in camera composition and camera planning. J. Blinn [Bli88] propose a vector algebra-based solution to pin two objects at given locations on the screen. Gleicher and Witkin [GW92] offer means to control 2D points directly on the display screen through differential manipulation techniques instead of controlling the camera parameters. Blinn's algebraic method has laid the groundwork for cinematographic idiom-based approaches, such as Christianson et al.'s compiler for the Declarative Camera Control Language (DCCL) [CAH*96] or He's et al.'s Virtual Cinematographer [HCS96]. These approaches however suffer from the point representation of the objects that doesn't capture the occlusion properties and thus avoids one of the main problems in camera planning, namely occlusion. Moreover the "point-like" representation cannot take into account the real geometry of the scene. Camera composition and camera planning can be viewed as constrained optimization problems in which the properties of the shot are expressed as numerical constraints on the camera variables (respectively the camera's path variables) and a broad range of solving procedures is available to compute solutions. The solvers differ in the way to manage over-constrained and under-constrained cases, in their complete or incomplete search capacities, in local minima management and possible optimization processes (generally finding the best solution w.r.t. an objective function). In their CAMDROID system for automated camera planning [DZ95], Drucker et al. use an existing numerical constraint solver package (CFSQP) and have to compile the shot specifications in the input script used by this library. Unfortunately, the solving process is sensitive to the initial configuration and is subject to local minima failures. Bares et al. use a partial constraint satisfaction system named CONSTRAINTCAM [BGL98] in order to provide alternate solutions when constraints cannot be completely satisfied. This solution is based on a limited subset of cinematographic properties (viewing angle, viewing distance and occlusion avoidance), which limits the procedure to small problems. P. Olivier et al.[OHPL99] propose a rich set of properties for composition purposes. The authors have developed the CAMPLAN system [HO00] that numerically solves the optimization problem via a metaheuristic search (genetic algorithms) method. The main shortcomings of this purely optimization-based technique is that the CAMPLAN genetic algorithm produces solutions in widely varying amounts of time and is also subject to the initial population of solutions. Gooch et al.[GRMS01] also use optimization procedures in order to produce images fitting artistic composition criterion like rules of thirds and fifths and the notion of canonical viewpoint (viewpoint leading to better identification of an object). The CSP (Constraint Satisfaction Problem) framework has proven to succeed in some camera composition and motion planning approaches. In [BMBT00] Bares et al. propose a heuristic-based complete search algorithm. The process is applied inside promising 3D areas computed through simple geometric intersections. Although efficient, the approach avoids problems related to multiple solutions. Unlike Bares et al. who utilize partial constraint satisfaction through cost functions, Jardillier & Languénou use pure interval methods in The Virtual Cameraman[JL98] to compute camera paths which yield sequences of images fulfilling temporally indexed image properties. This idea has been improved by Christie et al. in [CLG02]. Unfortunately, the benefits of interval-based techniques that guarantee the fulfillment of the properties during the whole sequence are counterbalanced by the computational effort required, and the absence of any mechanism for constraint relaxation. However, whenever the method fails, the user has a guarantee that there are no solutions to his problem (due to completeness of interval-based approaches). Some approaches combine constraints and optimization. One of these was presented by J. Pickering [Pic02] and is closely related to our work. In this solving method, the constraints are used to create feasible regions of space that will serve as bounds for an optimization procedure. The search space is subdivided by a "shadow-volumes" algorithm based on the properties of the image, and the feasible regions are then discretized and stored in an octree structure. Each node of the octree is then used as a starting point for a genetic algorithm that tries to find a solution to the problem. The main shortcomings of this approach lay in the computational effort required to create the octree, and the fact that multiple distinct solutions are ignored. Most of the methods we mention rely upon optimization processes to compute satisfactory camera placements and therefore lead to a unique solution closely related to the objective function. However, the description of a cinematic shot can possibly yield different visual solutions. Therefore, in computing the set of semantically distinct solutions w.r.t. cinematographic properties, one provides the user meaningful results. Figure 1 presents a top view of a simple scene containing three objects A, B and C. Whenever the user describes a shot in which he constrains A and B respectively to lay on the left and on the right of the screen, it clearly yields three possible classes of camera configurations: area (1) object C is on the left of A and B on the screen, (2) object C is between A and B, and (3) object C is on the right of A and B. Moreover, when considering possible occlusions, two classes can be added (4) A occludes C and (5) B occludes C. In such cases, classical optimization and incomplete CSP-based approaches fail in that a unique solution [DZ95,OHPL99,Pic02], or a reduced subset [JL98,CLG02] of solutions is proposed, whereas all classes of solutions should be equally considered. These approaches actually lose the semantics of the problem while relying upon pure numerical approaches. Once a solution is computed, no further information on its characteristics or differences with other possible solutions is provided. Certainly, any classification process can be provided afterhand, but requires an important computational effort. Possible distinct areas for viewing a couple of objects A and B (resp. on the left and right of the screen) w.r.t. to a third object C. In this paper, we propose to integrate a semantic dimension in the solving and interaction processes to assist the user in his camera placement tasks. We follow a threefold declarative approach: describe the desired solution with a set of cinematographic-based properties, compute distinct classes of solutions satisfying the description with related cinematographic properties, explore and interact with the classes of possible solutions. In the description phase, as in previous approaches [OHPL99,HO00,JL98,DZ95], a high-level grammar is offered including composition properties (framing objects on screen surface, relative object orientation and size) and shot properties (close shot, establishing shot, low and high angle). The geometry of the scene (locations and orientation of objects) is considered as an input provided by the user. In the computational phase, two processes are combined. The first process partitions the search space according to cinematographic properties (e.g. area such that A occludes B on the screen) and builds the intersection of the space partitions. The second process computes a nice representative of each possible class of solutions via a continuous domain implementation of a local search metaheuristic algorithm. Finally, in a third phase, the user navigates in the possible solution sets and interacts with the semantic information provided in each area. This paper concentrates on the first two phases and offers solid foundations for high-level interactions with the user. This paper is organized as follows: Section 2 introduces our semantic space partitioning approach to virtual camera composition, Section 3 details the numerical solving process. The exploitation of the semantic volumes is presented in Section 4 and relevant results are then presented in Section 5. Finally Section 6 discusses future research directions and concludes. Our approach to virtual camera compostion (VCC) is based on the primary idea of Binary Space Partition (BSP) and can be considered as an extension of visual aspects[KvD79] and closely related works such as viewpoint space partitioning[PD90] in the field of object recognition. The idea behind visual aspects is to gather all the viewpoints of a single polyhedron that share similar topological characteristics on the image. A change of appearance of the polyhedron with changing viewpoint, gives rise to boundaries in the search space. Computing all the boundaries enables the construction of regions of constant aspect, namely viewpoint space partitions. In this paper, we propose an extension of viewpoint space partitions to multiple objects and replace the topological characteristics of a polyhedron by cinematographic properties such as occlusions, relative viewing angles, distance shots and relative object locations. We introduce the notion of semantic volume as a volume of possible camera locations that give rise to qualitatively equivalent shots w.r.t. to cinematographic properties, i.e. semantically equivalent shots. Each volume is characterized by a set of semantic tags issued from film grammar [Ari76] and each tag is associated to a satisfied property in the volume. Tags are either related to a single object such as viewing angle, or to a couple of objects such as occlusion and relative image location. Therefore, the entire space of possible camera locations is thoroughly partitioned for each object, and each couple of objects. We then derive from the user's description the subset of volumes to be considered and intersect them. This process leads to a set of non-connected regions of which each represents a different class of solutions in terms of visual aspect. The computation of semantic volumes can be formalized as follows. We define the geometric filtering operator Gƒ that inputs a property p and provides a semantic volume sv, which is defined by sv=〈S, V〉 where S is a conjunction of semantic tags (e.g. LeftOf(A) ΛMediumShotOn(B) ΛOccludes(A,B) Λ…) and V is a subset of possible camera locations in ℜ3. Every camera location inside sv possibly satisfies the property p, whereas every camera location outside sv certainly violates p. Possible camera orientations are to be further computed by the numerical process (see Section 3). The filtering operator is complete in that it does not lose any correct camera locations. The operator Gƒ computes (possibly) non-connected volumes by pruning the space of most of the inconsistent camera locations w.r.t. p and associates a semantic tag to the resulting volumes. For example, the user description "A occludes B" (see Fig. 5) leads to a cone-shaped volume V coupled with the semantic tag Occlusion (B,A). Occlusion cones computation (both partial and total occlusions). The following subsections present how the filtering operator Gf provides the semantic volumes related to each property. Below, we consider that the camera's roll degree of freedom (i.e. around the look-at vector) is restricted to interval , and the tilt angle is confined in (no upside-down cameras). Camera compositions seldom break these conventions. The projection property is based on the notion of scale shots in cinematography (cf.Fig 2). It allows the artist to specify a viewing shot for an object. There are basically six different kinds of shots : the Extreme Close-Up, the Close-Up, the Medium Close-Up, the Medium Long Shot (or Plan Américain), the Long Shot, the Extreme Long Shot. Related semantic tags reflect all six kinds of shots (ExtremeCloseUp(Object) to ExtremeLongShot(Object)). The six distances between the camera and a character according to Arijon [Ari76]. The underlying semantic volume is computed given the position of an object and a cinematographic scale shot specified by the user. One can deduce an optimal size corresponding to each scale shot presented in Figure 2. The object's bounding sphere and desired area in the frame are used to determine the range of camera distances. In order to add some flexibility to the solving system, we compute an interval range related to this optimal value by subtracting and adding an epsilon to it (see [BMBT00]). The minimum and maximum bounds of the interval correspond to two distances defining the inner and outer radiuses of an hollow sphere that includes the set of consistent positions for the camera (cf.Fig. 3). Distance is trivially computed by the following equation: Semantic volumes leading to characteristics shots. The orientation property lets the virtual cinematographer specify the viewing angle required to shoot an object or a character. A common set of 8 viewing angles is offered (e.g. relative to object A, there are IsLeftProfileOf(A), IsRightProfileOf(A), IsInFrontOf(A), IsInBackOf(A), IsThreeQuaterFrontLeft(A) up to IsThreeQuaterBackRight(A)) and each can be composed with high and low relative angles IsHighAngle(A) and IsLowAngle(A). Computing orientation semantic volumes consists in building a prism-shaped volume of possible camera locations, w.r.t. the vector to consider (front, back, left, …). Once again, a variation with the optimal orientation is accepted in order to avoid being too restrictive (cf.Fig. 4). Four common relative viewing angles and related semantic tags. The occlusion property gives the user the opportunity to specify some visibility constraints between two objects of the scene. The cinematographer can characterize a total occlusion of an object by another, a partial occlusion or an absence of occlusion between two objects. A partial occlusion occurs when a part of an object's projection overlaps some part of the second object's projection. The semantic volumes induced by an occlusion property are computed given characteristic cones defined with respect to the positions of the two objects involved in the occlusionproperty [DDP02]. The inside bounds of the cones define the volumes of partial occlusion. The outside bounds of the cones define the volumes where no possible occlusion can occur (cf.Fig. 5). The framing property allows the virtual cinematographer to constrain an object in a given frame inside, partially inside, or outside the screen space. This expressive property defines the relative locations and sizes of objects on the screen. It constrains altogether the distance shot and the camera locations and orientations. Moreover, total or partial occlusions can be derived from overlapping frames, and conversely non-occlusions can be derived from the absence of overlap (cf.Fig. 6and 8). A non-overlapping description constraining three objects. An overlapping description constraining three objects with an overlap. The computation and characterization of the semantic volumes related to this property depend on the number of framing properties defined by the user and their relative location on the screen. We two (1) whenever a single frame is the framing property will like a projection the size of the frame and the size of the object in the scene lead to the computation of a volume that limits the all possible camera locations around the object are possible w.r.t. this (2) the user has specified two or more a is to be for every couple of objects in order to The of all the distinct frame is not provided we propose a two leading to different semantic volumes considering overlapping and non-overlapping 8 and that the semantic volumes computed at the of the framing will the between the on the screen (e.g. object A is left of but not the locations of the objects in the the numerical process will compute camera orientations to locations in the Figure 6 a user input related to a non-overlapping Figure provides a 2D of the scene with locations of objects A, B and C. For the representation is two but all the computation occurs in The first semantic volume is as 1 A and and with possible camera location in this area can lead to a shot in which the object A on the left of B. as 1 and 2 not overlap on the screen, no occlusion should occur between A and B. 2 and 3 are by computing the occlusion cones [DDP02]. tags are where B A and when A B. the area of Figure for the possible camera locations satisfying the user's description 2D of the semantic volumes related to couple , when respectively on the left and right of the screen. This process is for every couple of objects in the screen, Figure 8 an overlapping the overlapping is partial and areas and are not considered as possible camera locations and are with or the area between A and B is not considered as possible 3 does not provide any overlapping configuration and is with in the areas containing possible camera locations are at the left and right of Figure distinct volumes are the camera is in the right solution B will than A in the shot, whereas it is to the left, A will than B. 2D of semantic volumes related to overlapping on objects A and B the further semantic partitioning and characterization is offered by considering and The of the whole set of distinct frame follow the process. positioning properties allow to specify some relative positions between objects on the screen, one can that one wants to see the object A on the right of B, B, These relative placements describe the of objects being to within a visual composition as a conjunction of properties is by a intersection of the semantic volumes. a set of properties provided by the user results in the following computation : The intersection process lead to an a unique volume or to a set of non-connected volumes. In order to the intersection of 3D we propose to rely on of the semantic volumes than pure geometric Each is the of a field the function. For each point of the the of the are A 3D volume is when a value such that provides with the following (1) of volumes by the (2) in that we avoid or (3) to a point in the via a and simple Our implementation relies upon the that provides and means to create by defining for each object and with the bounding the whole 3D scene. For example, the possible camera locations the scale shots distances between the camera and an object (cf.Fig. are via the volumes the occlusion property are defined as cones the objects (see Fig. thus the regions of total or partial The to the use of in the computational cost of our approach requires a unique at the in order to determine the number of non-connected The cost of the is by the number of and for our the computation of the non-connected with The of the space partitioning approach consists in a semantic volume V containing possible camera locations. the numerical computes a nice representative of each volume in This consists in in V a consistent camera configuration orientation and that satisfies the framing properties and that each property corresponding to a semantic tag of S a cost function). The problem therefore to determine a of variables such that : where for the cost associated to property and where is the framing property given by the user. optimization and constrained optimization techniques a objective (e.g. In order to manage algebraic constraints such as and algebraic constraints such as we rely on The underlying of the framework can be in 1 presents our The algorithm relies on the of the search space a semantic starting from an initial and the around the current The a set of in V within a around the current It introduces the notions of and The allows on promising regions of the search space by the size the being in local no has been for a while a of the initial is to allow a of the search space. The procedure is by the maximum number of and and the number of at each of the function. each the best in terms of constraint satisfaction and cost the new current the user's description not lead to a solution framing is to respect as objects have been in the 3D we propose to integrate the satisfaction of the framing constraints in the cost function. Therefore, is given by : The cost to the respect of a property considering the orientation property of for example, the best are vector and such that the camera's orientation is to This local search technique be used to compute solutions of camera composition problems to Olivier et al.'s However, the process the solver in promising regions and the of the associated to each property. The constraint is by a simple of the related to a semantic the configuration in V and conversely is outside The main of this paper is to offer a semantic for and with the volumes. We two possible on the computed semantic volumes and on the whole 3D scene the computed volumes. For a given each computed volume provided by the geometric solver some related to the satisfaction of the properties. the characterization of each distinct volume can be semantically according to properties the user has not For example, a semantic volume sv is characterized by the relative location some further characterization of sv can be computed by considering the orientation properties related to A and B (e.g. a the user can two semantically volumes and the description and for the differences between them. This to compute all tags in and in that not each object and couple of objects in the scene their semantic it is to on a computed volume sv any possible object or property. The to computing a new geometric intersection and the number of computed by For example, in 2 of Section one can there a possible camera location such that A, B and C can be viewed from the relative angle : where sv is the computed semantic volume and Gf the geometric that computes the semantic volume related to the property is the is all possible volumes the a computed volume w.r.t. a property a similar For example, the of a volume sv into all the possible shot distances relative to an object A can be expressed as : The set of properties in Section 2 is provided via a as an extension to the It is for the and the of the solutions given to the user. Moreover it the into existing by the of and 2 the related to most properties. The framing property as parameters the of the left and top right of the frame containing the object. less than or more than 1 frame objects of the screen. The orientation property requires an object and a viewing angle presented in Section The projection property an object and a shot class 2). Occlusion properties associates with and the occlusion or our the classical shot two and then a framing shot with possible classes of solutions. The classical shot is when a between two or more and consists in the camera behind one while framing the Figure the related semantic volume and a result is presented in Figure The shot is given by the following declarative script w.r.t. and view of the search space computed by the shot and of the volume A result of the The geometry of the scene is composed of objects (see Fig. The user objects A, B and C respectively in the left, and right of the screen, and constrains and to to the screen any occlusion. Some results are presented in Figure three shots the users description and different classes of solutions. view of the shot and related to possible camera locations with further semantic information related to orientations of A , B and C shots associated to the framing 3 presents the time during the and numerical in computation of one representative of each semantic volume. is directly related to the of the are and 1 and 2 respectively and is the time in the local search with and provide a Although the total time to compute all important time representative is around which is for interaction purposes. The semantic approach offers the following the cinematographic properties provide semantic volumes containing the possible solutions of the problem through a geometric process that areas of the The computation of the boundaries relies on and avoids volume Whenever the intersection process leads to an there is a guarantee of in the user's The numerical process offers a representative of each volume at low computational The use of as object boundaries in the occlusion computation can be present the of being for bounding objects that are one or for This lead to regions of the search space that possibly lead to correct shots during occlusions and results can be through but to integrate in our can be by computing volumes provided an representation of such volumes is The computational cost of our approach is related to the number of objects in the scene and to the user's Most time is in the computation of the number of non-connected volumes which is related to the of the function. We to this process by computing the intersection of of each semantic volume and then the on this Finally, the approach to is our main objective being to characterize possible camera positions for a of time the computation of camera A extension is volumes. the extension is not and the main in with the time criterion in each volume. in this paper we have presented an approach to virtual camera composition that classes of distinct provides means to characterize and computes the notion of visual we the notion of Semantic as a set of possible camera locations that share a set of cinematographic results the of our approach and in and to virtual camera
Marc Christie, Jean-Marie Normand
Comput. Graph. Forum2