VLDB 2026 Research / reviewers in the wild / expert
Marc Christie
dblp:63/5800
· DBLP profile ↗
57ranked-venue papers
3as first author
22since 2021 · last 2026
0000-0001-6080-8026ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 46 · 2 first-author · 18 since 2021Artificial intelligence and machine learning · 22 · 1 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 8 · 5 since 2021Systems, architecture and hardware · 3 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorComputer networks · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GTAvatar: Bridging Gaussian Splatting and Texture Mapping for Relightable and Editable Gaussian AvatarsabstractAbstract Recent advancements in Gaussian Splatting have enabled increasingly accurate reconstruction of photorealistic head avatars, opening the door to numerous applications in visual effects, videoconferencing, and virtual reality. This, however, comes with the lack of intuitive editability offered by traditional triangle mesh‐based methods. In contrast, we propose a method that combines the accuracy and fidelity of 2D Gaussian Splatting with the intuitiveness of UV texture mapping. By embedding each canonical Gaussian primitive's local frame into a patch in the UV space of a template mesh in a computationally efficient manner, we reconstruct continuous editable material head textures from a single monocular video on a conventional UV domain. Furthermore, we leverage an efficient physically based reflectance model to enable relighting and editing of these intrinsic material maps. Through extensive comparisons with state‐of‐the‐art methods, we demonstrate the accuracy of our reconstructions, the quality of our relighting results, and the ability to provide intuitive controls for modifying an avatar's appearance and geometry via texture mapping without additional optimization. Kelian Baert, Mae Younes, Francois Bourel, Marc Christie, Adnane Boukhayma |
Comput. Graph. Forum | 4 |
| 2025 | AKiRa: Augmentation Kit on Rays for Optical Video GenerationabstractRecent advances in text-conditioned video diffusion have greatly improved video quality. However, these methods offer limited or sometimes no control to users on camera aspects, including dynamic camera motion, zoom, distorted lens and focus shifts. These motion and optical aspects are crucial for adding controllability and cinematic elements to generation frameworks, ultimately resulting in visual content that draws focus, enhances mood, and guides emotions according to filmmakers’ controls. In this paper, we aim to close the gap between controllable video generation and camera optics. To achieve this, we propose AKiRa (Augmentation Kit on Rays), a novel augmentation framework that builds and trains a camera adapter with a complex camera model over an existing video generation backbone. It enables fine-tuned control over camera motion as well as complex optical parameters (focal length, distortion, aperture) to achieve cinematic effects such as zoom, fish-eye effect, and bokeh. Extensive experiments demonstrate AKiRa’s effectiveness in combining and composing camera optics while outperforming all state-of-the-art methods. This work sets a new landmark in controlled and optically enhanced video generation, paving the way for future optical video generation methods. Xi Wang 0024, Robin Courant, Marc Christie, Vicky Kalogeiton |
CVPR | 3 |
| 2025 | Foreword to special section on expressive media
Chiara Eva Catalano, Amal Dev Parakkat, Marc Christie |
Comput. Graph. | 3 |
| 2024 | E.T. the Exceptional Trajectories: Text-to-Camera-Trajectory Generation with Character Awareness
Robin Courant, Nicolas Dufour, Xi Wang 0024, Marc Christie, Vicky Kalogeiton |
ECCV (4) | 4 |
| 2024 | SPARK: Self-supervised Personalized Real-time Monocular Face CaptureabstractFeedforward monocular face capture methods seek to reconstruct posed faces from a single image of a person. Current state of the art approaches have the ability to regress parametric 3D face models in real-time across a wide range of identities, lighting conditions and poses by leveraging large image datasets of human faces. These methods however suffer from clear limitations in that the underlying parametric face model only provides a coarse estimation of the face shape, thereby limiting their practical applicability in tasks that require precise 3D reconstruction (aging, face swapping, digital make-up, ...). In this paper, we propose a method for high-precision 3D face capture taking advantage of a collection of unconstrained videos of a subject as prior information. Our proposal builds on a two stage approach. We start with the reconstruction of a detailed 3D face avatar of the person, capturing both precise geometry and appearance from a collection of videos. We then use the encoder from a pre-trained monocular face reconstruction method, substituting its decoder with our personalized model, and proceed with transfer learning on the video collection. Using our pre-estimated image formation model, we obtain a more precise self-supervision objective, enabling improved expression and pose alignment. This results in a trained encoder capable of efficiently regressing pose and expression parameters in real-time from previously unseen images, which combined with our personalized geometry model yields more accurate and high fidelity mesh inference. Through extensive qualitative and quantitative evaluation, we showcase the superiority of our final model as compared to state-of-the-art baselines, and demonstrate its generalization ability to unseen pose, expression and lighting. Kelian Baert, Shrisha Bharadwaj, Fabien Castan, Benoit Maujean, Marc Christie, Victoria Fernández Abrevaya, Adnane Boukhayma |
SIGGRAPH Asia | 5 |
| 2024 | Cinematographic Camera Diffusion ModelabstractAbstract Designing effective camera trajectories in virtual 3D environments is a challenging task even for experienced animators. Despite an elaborate film grammar, forged through years of experience, that enables the specification of camera motions through cinematographic properties (framing, shots sizes, angles, motions), there are endless possibilities in deciding how to place and move cameras with characters. Dealing with these possibilities is part of the complexity of the problem. While numerous techniques have been proposed in the literature (optimization‐based solving, encoding of empirical rules, learning from real examples,…), the results either lack variety or ease of control. In this paper, we propose a cinematographic camera diffusion model using a transformer‐based architecture to handle temporality and exploit the stochasticity of diffusion models to generate diverse and qualitative trajectories conditioned by high‐level textual descriptions. We extend the work by integrating keyframing constraints and the ability to blend naturally between motions using latent interpolation, in a way to augment the degree of control of the designers. We demonstrate the strengths of this text‐to‐camera motion approach through qualitative and quantitative experiments and gather feedback from professional artists. The code and data are available at https://github.com/jianghd1996/Camera-control . Hongda Jiang, Xi Wang 0024, Marc Christie, Libin Liu 0002, Baoquan Chen |
Comput. Graph. Forum | 3 |
| 2024 | Real-Time Multi-Map Saliency-Driven Gaze Behavior for Non-Conversational CharactersabstractGaze behavior of virtual characters in video games and virtual reality experiences is a key factor of realism and immersion. Indeed, gaze plays many roles when interacting with the environment; not only does it indicate what characters are looking at, but it also plays an important role in verbal and non-verbal behaviors and in making virtual characters alive. Automated computing of gaze behaviors is however a challenging problem, and to date none of the existing methods are capable of producing close-to-real results in an interactive context. We therefore propose a novel method that leverages recent advances in several distinct areas related to visual saliency, attention mechanisms, saccadic behavior modelling, and head-gaze animation techniques. Our approach articulates these advances to converge on a multi-map saliency-driven model which offers real-time realistic gaze behaviors for non-conversational characters, together with additional user-control over customizable features to compose a wide variety of results. We first evaluate the benefits of our approach through an objective evaluation that confronts our gaze simulation with ground truth data using an eye-tracking dataset specifically acquired for this purpose. We then rely on subjective evaluation to measure the level of realism of gaze animations generated by our method, in comparison with gaze animations captured from real actors. Our results show that our method generates gaze behaviors that cannot be distinguished from captured gaze animations. Overall, we believe that these results will open the way for more natural and intuitive design of realistic and coherent gaze animations for real-time applications. Ific Goudé, Alexandre Bruckert, Anne-Hélène Olivier, Julien Pettré, Rémi Cozot, Kadi Bouatouch, Marc Christie, Ludovic Hoyet |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2024 | With or Without You: Effect of Contextual and Responsive Crowds on VR-based Crowd Motion CaptureabstractWhile data is vital to better understand and model interactions within human crowds, capturing real crowd motions is extremely challenging. Virtual Reality (VR) demonstrated its potential to help, by immersing users into either simulated virtual crowds based on autonomous agents, or within motion-capture-based crowds. In the latter case, users' own captured motion can be used to progressively extend the size of the crowd, a paradigm called Record-and-Replay (2R). However, both approaches demonstrated several limitations which impact the quality of the acquired crowd data. In this paper, we propose the new concept of contextual crowds to leverage both crowd simulation and the 2R paradigm towards more consistent crowd data. We evaluate two different strategies to implement it, namely a Replace-Record-Replay (3R) paradigm where users are initially immersed into a simulated crowd whose agents are successively replaced by the user's captured-data, and a Replace-Record-Replay-Responsive (4R) paradigm where the pre-recorded agents are additionally endowed with responsive capabilities. These two paradigms are evaluated through two real-world-based scenarios replicated in VR. Our results suggest that the behaviors observed in VR users with surrounding agents from the beginning of the recording process are made much more natural, enabling 3R or 4R paradigms to improve the consistency of captured crowd datasets. Tairan Yin, Ludovic Hoyet, Marc Christie, Marie-Paule Cani, Julien Pettré |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2023 | The Stare-in-the-Crowd Effect When Navigating a Crowd in Virtual RealityabstractNonverbal communication is paramount in daily life, as well as in populated virtual reality (VR) environments. In this paper, we focused on gaze behaviour, which is key to initiate and drive social interactions. Previous work on photographs and on virtual agents showed the importance of gaze, even in the presence of multiple stimuli, by demonstrating the stare-in-the-crowd effect: humans detect faster and observe gazes directed towards them longer than the averted ones. While previous studies focused on static scenarios, which fail in representing the complexity of real-life social interactions, we propose to explore the stare-in-the-crowd effect in dynamic situations. To this end, we designed a within-subject experiment where 21 users navigated a virtual street through an idle or moving crowd of virtual agents. Agents’ gaze was manipulated to display averted, directed, or shifting gaze. We analysed the user’s gaze (fixations, dwell time) and locomotor behaviours (path decisions, proximity to agents) as well as their social anxiety. Results showed that the stare-in-the-crowd effect is preserved when walking through both types of crowd, and that social anxiety decreases gaze interaction time and affects proximity behaviours in case of agents with directed gazes. However, virtual agents’ gaze did not elicit significant changes on users’ locomotion. These findings highlight the importance of considering virtual agents’ gaze when creating VR environments, and open future work perspectives to better understand factors that would strengthen or decrease this effect at gaze and locomotor levels. Pierre Raimbaud, Alberto Jovane, Katja Zibrek, Claudio Pacchierotti, Marc Christie, Ludovic Hoyet, Julien Pettré, Anne-Hélène Olivier |
SAP | 5 |
| 2023 | JAWS: Just A Wild Shot for Cinematic Transfer in Neural Radiance FieldsabstractThis paper presents JAWS, an optimization-driven approach that achieves the robust transfer of visual cinematic features from a reference in-the-wild video clip to a newly generated clip. To this end, we rely on an implicit-neural-representation (INR) in a way to compute a clip that shares the same cinematic features as the reference clip. We propose a general formulation of a camera optimization problem in an INR that computes extrinsic and intrinsic camera parameters as well as timing. By leveraging the differentiability of neural representations, we can back-propagate our designed cinematic losses measured on proxy estimators through a NeRF network to the proposed cinematic parameters directly. We also introduce specific enhancements such as guidance maps to improve the overall quality and efficiency. Results display the capacity of our system to replicate well known camera sequences from movies, adapting the framing, camera parameters and timing of the generated video clip to maximize the similarity with the reference clip. Robin Courant, Jinglei Shi, Éric Marchand, Marc Christie |
CVPR | 5 |
| 2023 | Real-time Computational Cinematographic Editing for Broadcasting of Volumetric-captured events: an Application to Ultimate FightingabstractThe capacity to capture and broadcast sports events with close to real-time volumetric reconstruction techniques opens exciting perspectives in how audiences can consume and interact with these contents. In this work, we propose the design of an real-time cinematography system that is capable of generating qualitative framing and editing of volumetric-captured content, by mimicking real broadcast footage. To illustrate our approach, we focus on the specific problem of cinematography for ring-based events such as Ultimate Fighting Championships (UFC). We start by extracting statistical features from hours of real footage to understand the specific framing and cutting behaviors of real broadcast directors. We then exploit these features in a real-time editing system, not only to replicate the behaviors in existing broadcasting, but also to generalize to novel camera layouts. We demonstrate our approach on the volumetric reconstruction of UFC fights, compare our results with baseline methods, and report different qualitative and quantitative evaluations. Francois Bourel, Xi Wang 0024, Ervin Teng, Valerio Ortenzi, Adam Myhill, Marc Christie |
MIG | 6 |
| 2023 | Warping character animations using visual motion features
Alberto Jovane, Pierre Raimbaud, Katja Zibrek, Claudio Pacchierotti, Marc Christie, Ludovic Hoyet, Anne-Hélène Olivier, Julien Pettré |
Comput. Graph. | 5 |
| 2023 | Contact-conditioned hand-held object reconstruction from single-view images
Yang Li 0041, Adnane Boukhayma, Changbo Wang, Marc Christie |
Comput. Graph. | 5 |
| 2022 | A new framework for the evaluation of locomotive motion datasets through motion matching techniquesabstractAnalyzing motion data is a critical step when building meaningful locomotive motion datasets. This can be done by labeling motion capture data and inspecting it, through a planned motion capture session or by carefully selecting locomotion clips from a public dataset. These analyses, however, have no clear definition of coverage, making it harder to diagnose when something goes wrong, such as a virtual character not being able to perform an action or not moving at a given speed. This issue is compounded by the large amount of information present in motion capture data, which poses a challenge when trying to interpret it. This work provides a visualization and an optimization method to streamline the process of crafting locomotive motion datasets. It provides a more grounded approach towards locomotive motion analysis by calculating different quality metrics, such as: demarcating coverage in terms of both linear and angular speeds, frame use frequency in each animation clip, deviation from the planned path, number of transitions, number of used vs. unused animations and transition cost. Vicenzo Abichequer Sangalli, Ludovic Hoyet, Marc Christie, Julien Pettré |
MIG | 3 |
| 2022 | The Stare-in-the-Crowd Effect in Virtual RealityabstractNonverbal cues are paramount in real-world interactions. Among these cues, gaze has received much attention in the literature. In particular, previous work has shown a search asymmetry between directed and averted gaze towards the observer using photographic stimuli, with faster detection and longer fixation towards directed gaze by the observer. This is known as the stare-in-the-crowd effect. In this study, we investigate whether stare-in-the crowd effect is preserved in Virtual Reality (VR). To this end, we designed a within-subject experiment where 30 human users were immersed in a virtual environment in front of an audience of 11 virtual agents following 4 different gaze behaviours. We analysed the user’s gaze behaviour when observing the audience, computing fixations and dwell time. We also collected the users’ social anxiety score using a post-experiment questionnaire to control for some potential influencing factors. Results show that the stare-in-the-crowd effect is preserved in VR, as demonstrated by the significant differences between gaze behaviours, similarly to what was found in previous studies using photographic stimuli. Additionally, we found a negative correlation between dwell time towards directed gazes and users’ social anxiety scores. Such results are encouraging for the development of expressive and reactive virtual humans, which can be animated to express natural interactive behaviour. Pierre Raimbaud, Alberto Jovane, Katja Zibrek, Claudio Pacchierotti, Marc Christie, Ludovic Hoyet, Julien Pettré, Anne-Hélène Olivier |
VR | 5 |
| 2022 | The One-Man-Crowd: Single User Generation of Crowd Motions Using Virtual RealityabstractCrowd motion data is fundamental for understanding and simulating realistic crowd behaviours. Such data is usually collected through controlled experiments to ensure that both desired individual interactions and collective behaviours can be observed. It is however scarce, due to ethical concerns and logistical difficulties involved in its gathering, and only covers a few typical crowd scenarios. In this work, we propose and evaluate a novel Virtual Reality based approach lifting the limitations of real-world experiments for the acquisition of crowd motion data. Our approach immerses a single user in virtual scenarios where he/she successively acts each crowd member. By recording the past trajectories and body movements of the user, and displaying them on virtual characters, the user progressively builds the overall crowd behaviour by him/herself. We validate the feasibility of our approach by replicating three real experiments, and compare both the resulting emergent phenomena and the individual interactions to existing real datasets. Our results suggest that realistic collective behaviours can naturally emerge from virtual crowd data generated using our approach, even though the variety in behaviours is lower than in real situations. These results provide valuable insights to the building of virtual crowd experiences, and reveal key directions for further improvements. Tairan Yin, Ludovic Hoyet, Marc Christie, Marie-Paule Cani, Julien Pettré |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2021 | Real-Time Cinematic Tracking of Targets in Dynamic EnvironmentsabstractTracking in a cinematic way a moving target inside a 3D dynamic environment remains a challenging problem. This requires to simultaneously ensure a low computational cost, a good degree of reactivity and a high cinematic quality despite sudden changes. In this paper, we draw on the idea of Motion-Predictive Control to propose an efficient real-time camera tracking technique which ensures these properties. Our approach relies on the predicted motion of a target to create and evaluate a very large number of camera motions using hardware ray casting. Our evaluation of camera motions includes a range of cinematic properties such as distance to target, visibility, collision, smoothness and jitter. Experiments are conducted to display the benefits of the approach with relation to prior work. Ludovic Burg, Christophe Lino, Marc Christie |
Graphics Interface | 3 |
| 2021 | Structures in Tropes Networks: Toward a Formal Story Grammar
Jean-Peic Chou, Marc Christie |
ICCC | 2 |
| 2021 | TT-SLAM: Dense Monocular SLAM for Planar EnvironmentsabstractThis paper proposes a novel visual SLAM method with dense planar reconstruction using a monocular camera: TT-SLAM. The method exploits planar template-based trackers (TT) to compute camera poses and reconstructs a multi-planar scene representation. Multiple homographies are estimated simultaneously by clustering a set of template trackers supported by superpixelized regions. Compared to RANSAC-based multiple homographies method [1], data association and keyframe selection issues are handled by the continuous nature of template trackers. A non-linear optimization process is applied to all the homographies to improve the precision in pose estimation. Experiments show that the proposed method outperforms RANSAC-based multiple homographies method [1] as well as other dense method SLAM techniques such as LSD-SLAM or DPPTAM, and competes with keypoint-based techniques like ORB-SLAM while providing dense planar reconstructions of the environment. Marc Christie, Éric Marchand |
ICRA | 2 |
| 2021 | Reactive Virtual Agents: A Viewpoint-Driven Approach for Bodily Nonverbal CommunicationabstractNon-verbal communication body cues are paramount to interact. In this preliminary work, we explore ways to let Intelligent Virtual Agents (IVAs) simulating nonverbal communication capabilities. We propose an approach to control IVAs' reactive behaviour from the analysis of other agents' apparent motions, in a situation of "observed" IVAs that act and "observers" that react. For that, first a viewpoint-driven analysis of the observed agent's motion is done, and then a synthesis of this analysis induces the observers' reaction. Pierre Raimbaud, Alberto Jovane, Katja Zibrek, Claudio Pacchierotti, Marc Christie, Ludovic Hoyet, Julien Pettré, Anne-Hélène Olivier |
IVA | 5 |
| 2021 | Deep saliency models : The quest for the loss function
Alexandre Bruckert, Hamed Rezazadegan Tavakoli, Zhi Liu 0003, Marc Christie, Olivier Le Meur |
Neurocomputing | 4 |
| 2021 | Camera keyframing with style and controlabstractWe present a novel technique that enables 3D artists to synthesize camera motions in virtual environments following a camera style , while enforcing user-designed camera keyframes as constraints along the sequence. To solve this constrained motion in-betweening problem, we design and train a camera motion generator from a collection of temporal cinematic features (camera and actor motions) using a conditioning on target keyframes. We further condition the generator with a style code to control how to perform the interpolation between the keyframes. Style codes are generated by training a second network that encodes different camera behaviors in a compact latent space, the camera style space. Camera behaviors are defined as temporal correlations between actor features and camera motions and can be extracted from real or synthetic film clips. We further extend the system by incorporating a fine control of camera speed and direction via a hidden state mapping technique. We evaluate our method on two aspects: i) the capacity to synthesize style-aware camera trajectories with user defined keyframes; and ii) the capacity to ensure that in-between motions still comply with the reference camera style while satisfying the keyframe constraints. As a result, our system is the first style-aware keyframe in-betweening technique for camera control that balances style-driven automation with precise and interactive control of keyframes. Hongda Jiang, Marc Christie, Xi Wang 0024, Libin Liu 0002, Bin Wang 0021, Baoquan Chen |
ACM Trans. Graph. | 2 |
| 2020 | Relative Pose Estimation and Planar Reconstruction via Superpixel-Driven Multiple HomographiesabstractThis paper proposes a novel method to simultaneously perform relative camera pose estimation and planar reconstruction of a scene from two RGB images. We start by extracting and matching superpixel information from both images and rely on a novel multi-model RANSAC approach to estimate multiple homographies from superpixels and identify matching planes. Ambiguity issues when performing homography decomposition are handled by proposing a voting system to more reliably estimate relative camera pose and plane parameters. A non-linear optimization process is also proposed to perform bundle adjustment that exploits a joint representation of homographies and works both for image pairs and whole sequences of image (vSLAM). As a result, the approach provides a mean to perform a dense 3D plane reconstruction from two RGB images only without relying on RGB-D inputs or strong priors such as Manhattan assumptions, and can be extented to handle sequences of images. Our results compete with keypointbased techniques such as ORB-SLAM while providing a dense representation and are more precise than direct and semi-direct pose estimation techniques used in LSD-SLAM or DPPTAM. Marc Christie, Éric Marchand |
IROS | 2 |
| 2020 | Topology-aware Camera Control for Real-time ApplicationsabstractPlacing and moving virtual cameras in real-time 3D environments is a task that remains complex due to the many requirements which need to be satisfied simultaneously. Beyond the essential features of ensuring visibility and frame composition for one or multiple targets, an ideal camera system should provide designers with tools to create variations in camera placement and motions, and create shots which conform to aesthetic recommendations. In this paper, we propose a controllable process that will assist developers and artists in placing cinematographic cameras and camera paths throughout complex virtual environments, a task that was often manually performed until now. With no specification and no previous knowledge on the events, our tool exploits a topological analysis of the environment to capture the potential movements of the agents, highlight linearities and create an abstract skeletal representation of the environment. This representation is then exploited to automatically generate potentially relevant camera positions and trajectories organized in a graph representation with visibility information. At run-time, the system can then efficiently select appropriate cameras and trajectories according to artistic recommendations. We demonstrate the features of the proposed system with realistic game-like environments, highlighting the capacity to analyze a complex environment, generate relevant camera positions and camera tracks, and run efficiently with a range of different camera behaviours. Alberto Jovane, Amaury Louarn, Marc Christie |
MIG | 3 |
| 2020 | An interactive staging-and-shooting solver for virtual cinematographyabstractResearch in virtual cinematography often narrows the problem down to computing the optimal viewpoint for the camera to properly convey a scene’s content. In contrast we propose to address simultaneously the questions of placing cameras, lights, objects and actors in a virtual environment through a high-level specification. We build on a staging language and propose to extend it by defining complex temporal relationships between these entities. We solve such specifications by designing pruning operators which iteratively reduce the range of possible degrees of freedom for entities while satisfying temporal constraints. Our solver first decomposes the problem by analyzing the graph of relationships between entities and then solves an ordered sequence of sub-problems. Users have the possibility to manipulate the current result for fine-tuning purposes or to creatively explore ranges of solutions while maintaining the relationships. As a result, the proposed system is the first staging-and-shooting cinematography system which enables the specification and solving of spatio-temporal cinematic layouts. Amaury Louarn, Quentin Galvane, Fabrice Lamarche, Marc Christie |
MIG | 4 |
| 2020 | Real-time Anticipation of Occlusions for Automated Camera Control in Toric SpaceabstractAbstract Efficient visibility computation is a prominent requirement when designing automated camera control techniques for dynamic 3D environments; computer games, interactive storytelling or 3D media applications all need to track 3D entities while ensuring their visibility and delivering a smooth cinematic experience. Addressing this problem requires to sample a large set of potential camera positions and estimate visibility for each of them, which in practice is intractable despite the efficiency of ray‐casting techniques on recent platforms. In this work, we introduce a novel GPU‐rendering technique to efficiently compute occlusions of tracked targets in Toric Space coordinates – a parametric space designed for cinematic camera control. We then rely on this occlusion evaluation to derive an anticipation map predicting occlusions for a continuous set of cameras over a user‐defined time window. We finally design a camera motion strategy exploiting this anticipation map to minimize the occlusions of tracked entities over time. The key features of our approach are demonstrated through comparison with traditionally used ray‐casting on benchmark scenes, and through an integration in multiple game‐like 3D scenes with heavy, sparse and dense occluders. Ludovic Burg, Christophe Lino, Marc Christie |
Comput. Graph. Forum | 3 |
| 2020 | Example-driven virtual cinematography by learning camera behaviorsabstractDesigning a camera motion controller that has the capacity to move a virtual camera automatically in relation with contents of a 3D animation, in a cinematographic and principled way, is a complex and challenging task. Many cinematographic rules exist, yet practice shows there are significant stylistic variations in how these can be applied. In this paper, we propose an example-driven camera controller which can extract camera behaviors from an example film clip and re-apply the extracted behaviors to a 3D animation, through learning from a collection of camera motions. Our first technical contribution is the design of a low-dimensional cinematic feature space that captures the essence of a film's cinematic characteristics (camera angle and distance, screen composition and character configurations) and which is coupled with a neural network to automatically extract these cinematic characteristics from real film clips. Our second technical contribution is the design of a cascaded deep-learning architecture trained to (i) recognize a variety of camera motion behaviors from the extracted cinematic features, and (ii) predict the future motion of a virtual camera given a character 3D animation. We propose to rely on a Mixture of Experts (MoE) gating+prediction mechanism to ensure that distinct camera behaviors can be learned while ensuring generalization. We demonstrate the features of our approach through experiments that highlight (i) the quality of our cinematic feature extractor (ii) the capacity to learn a range of behaviors through the gating mechanism, and (iii) the ability to generate a variety of camera motions by applying different behaviors extracted from film clips. Such an example-driven approach offers a high level of controllability which opens new possibilities toward a deeper understanding of cinematographic style and enhanced possibilities in exploiting real film data in virtual environments. Hongda Jiang, Bin Wang 0069, Marc Christie, Baoquan Chen |
ACM Trans. Graph. | 4 |
| 2019 | Deep Learning For Inter-Observer Congruency PredictionabstractAccording to the literature regarding visual saliency, observers may exhibit considerable variations in their gaze behaviors. These variations are influenced by aspects such as cultural background, age or prior experiences, but also by features in the observed images. The dispersion between the gaze of different observers looking at the same image is commonly referred as inter-observer congruency (IOC). Predicting this congruence can be of great interest when it comes to study the visual perception of an image. In this paper, we introduce a new method based on deep learning techniques to predict the IOC of an image. This is achieved by first extracting features from an image through a deep convolutional network. We then show that using such features to train a model with a shallow network regression technique significantly improves the precision of the prediction over existing approaches. Alexandre Bruckert, Yat Hong Lam, Marc Christie, Olivier Le Meur |
ICIP | 3 |
| 2019 | VR as a Content Creation Tool for Movie PrevisualisationabstractCreatives in animation and film productions have forever been exploring the use of new means to prototype their visual sequences before realizing them, by relying on hand-drawn storyboards, physical mockups or more recently 3D modelling and animation tools. However these 3D tools are designed in mind for dedicated animators rather than creatives such as film directors or directors of photography and remain complex to control and master. In this paper we propose a VR authoring system which provides intuitive ways of crafting visual sequences, both for expert animators and expert creatives in the animation and film industry. The proposed system is designed to reflect the traditional process through (i) a storyboarding mode that enables rapid creation of annotated still images, (ii) a previsualisation mode that enables the animation of the characters, objects and cameras, and (iii) a technical mode that enables the placement and animation of complex camera rigs (such as cameras cranes) and light rigs. Our methodology strongly relies on the benefits of VR manipulations to re-think how content creation can be performed in this specific context, typically how to animate contents in space and time. As a result, the proposed system is complimentary to existing tools, and provides a seamless back-and-forth process between all stages of previsualisation. We evaluated the tool with professional users to gather experts' perspectives on the specific benefits of VR in 3D content creation. Quentin Galvane, I-Sheng Lin, Fernando Arqelaquet, Tsai-Yen Li, Marc Christie |
VR | 5 |
| 2018 | Multiple Layers of Contrasted Images for Robust Feature-Based Visual TrackingabstractFeature-based SLAM (Simultaneous Localization and Mapping) techniques rely on low-level contrast information extracted from images to detect and track keypoints. This process is known to be sensitive to changes in illumination of the environment that can lead to tracking failures. This paper proposes a multi-layered image representation (MLI) that computes and stores different contrast-enhanced versions of an original image. Keypoint detection is performed on each layer, yielding better robustness to light changes. An optimization technique is also proposed to compute the best contrast enhancements to apply in each layer. Results demonstrate the benefits of MLI when using the main keypoint detectors from ORB, SIFT or SURF, and shows significant improvement in SLAM robustness. Marc Christie, Éric Marchand |
ICIP | 2 |
| 2018 | Optimized Contrast Enhancements to Improve Robustness of Visual Tracking in a SLAM Relocalisation ContextabstractRobustness of indirect SLAM techniques to light changing conditions remains a central issue in the robotics community. With the change in the illumination of a scene, feature points are either not extracted properly due to low contrasts, or not matched due to large differences in descriptors. In this paper, we propose a multi-layered image representation (MLI) in which each layer holds a contrast enhanced version of the current image in the tracking process in order to improve detection and matching. We show how Mutual Information can be used to compute dynamic contrast enhancements on each layer. We demonstrate how this approach dramatically improves the robustness in dynamic light changing conditions on both synthetic and real environments compared to default ORB-SLAM. This work focalises on the specific case of SLAM relocalisation in which a first pass on a reference video constructs a map, and a second pass with a light changed condition relocalizes the camera in the map. Marc Christie, Éric Marchand |
IROS | 2 |
| 2018 | Automated staging for virtual cinematographyabstractWhile the topic of virtual cinematography has essentially focused on the problem of computing the best viewpoint in a virtual environment given a number of objects placed beforehand, the question of how to place the objects in the environment with relation to the camera (referred to as staging in the film industry) has received little attention. This paper first proposes a staging language for both characters and cameras that extends existing cinematography languages with multiple cameras and character staging. Second, the paper proposes techniques to operationalize and solve staging specifications given a 3D virtual environment. The novelty holds in the idea of exploring how to position the characters and the cameras simultaneously while maintaining a number of spatial relationships specific to cinematography. We demonstrate the relevance of our approach through a number of simple and complex examples. Amaury Louarn, Marc Christie, Fabrice Lamarche |
MIG | 2 |
| 2018 | Directing the Photography: Combining Cinematic Rules, Indirect Light Controls and Lighting-by-ExampleabstractAbstract The placement of lights in a 3D scene is a technical and artistic task that requires time and trained skills. Most 3D modelling tools only provide a direct control of light sources, through the manipulation of parameters such as size, location, flux (the perceived power of light) or opening angle (the light frustum). Approaches have been relying on automated or semi‐automated techniques to relieve users from such low‐level manipulations at the expense of an important computational cost. In this paper, guided by discussions with experts in scene and object lighting, we propose an indirect control of area light sources. We first formalize the classical 3‐point lighting design principle (key‐light, fill‐lights and back/rim‐lights) in a parametric model. Given a key‐light placed in the scene, we then provide a computational approach to (i) automatically compute the position and size of fill‐lights and back/rim‐lights by analyzing the geometry of 3D character, and (ii) automatically compute the flux and size of key, fill and back/rim lights, given a sample reference image in a computationally efficient way. Results demonstrate the benefits of the approach on the quick lighting of 3D characters, and further demonstrate the feasibility of interactive control of multiple lights through image features. Quentin Galvane, Christophe Lino, Marc Christie, Rémi Cozot |
Comput. Graph. Forum | 3 |
| 2018 | Directing Cinematographic DronesabstractQuadrotor drones equipped with high-quality cameras have rapidly raised as novel, cheap, and stable devices for filmmakers. While professional drone pilots can create aesthetically pleasing videos in short time, the smooth—and cinematographic—control of a camera drone remains challenging for most users, despite recent tools that either automate part of the process or enable the manual design of waypoints to create drone trajectories. This article moves a step further by offering high-level control of cinematographic drones for the specific task of framing dynamic targets. We propose techniques to automatically and interactively plan quadrotor drone motions in dynamic three-dimensional (3D) environments while satisfying both cinematographic and physical quadrotor constraints. We first propose the Drone Toric Space , a dedicated camera parameter space with embedded constraints, and derive some intuitive on-screen viewpoint manipulators. Second, we propose a dedicated path planning technique that ensures both that cinematographic properties can be enforced along the path and that the path is physically feasible by a quadrotor drone. At last, we build on the Drone Toric Space and the specific path planning technique to coordinate the motion of multiple drones around dynamic targets. A number of results demonstrate the interactive and automated capacities of our approaches on different use-cases. Quentin Galvane, Christophe Lino, Marc Christie, Julien Fleureau, Fabien Servant, François-Louis Tariolle, Philippe Guillotel |
ACM Trans. Graph. | 3 |
| 2018 | Creating and chaining camera moves for quadrotor videographyabstractCapturing aerial videos with a quadrotor-mounted camera is a challenging creative task, as it requires the simultaneous control of the quadrotor's motion and the mounted camera's orientation. Letting the drone follow a pre-planned trajectory is a much more appealing option, and recent research has proposed a number of tools designed to automate the generation of feasible camera motion plans; however, these tools typically require the user to specify and edit the camera path, for example by providing a complete and ordered sequence of key viewpoints. In this paper, we propose a higher level tool designed to enable even novice users to easily capture compelling aerial videos of large-scale outdoor scenes. Using a coarse 2.5D model of a scene, the user is only expected to specify starting and ending viewpoints and designate a set of landmarks, with or without a particular order. Our system automatically generates a diverse set of candidate local camera moves for observing each landmark, which are collision-free, smooth, and adapted to the shape of the landmark. These moves are guided by a landmark-centric view quality field, which combines visual interest and frame composition. An optimal global camera trajectory is then constructed that chains together a sequence of local camera moves, by choosing one move for each landmark and connecting them with suitable transition trajectories. This task is formulated and solved as an instance of the Set Traveling Salesman Problem. Ke Xie 0001, Shengqiu Huang, Dani Lischinski, Marc Christie, Kai Xu 0004, Minglun Gong, Daniel Cohen-Or, Hui Huang 0004 |
ACM Trans. Graph. | 5 |
| 2018 | Thinking Like a Director: Film Editing Patterns for Virtual Cinematographic StorytellingabstractThis article introduces Film Editing Patterns (FEP) , a language to formalize film editing practices and stylistic choices found in movies. FEP constructs are constraints, expressed over one or more shots from a movie sequence, that characterize changes in cinematographic visual properties, such as shot sizes, camera angles, or layout of actors on the screen. We present the vocabulary of the FEP language, introduce its usage in analyzing styles from annotated film data, and describe how it can support users in the creative design of film sequences in 3D. More specifically, (i) we define the FEP language, (ii) we present an application to craft filmic sequences from 3D animated scenes that uses FEPs as a high level mean to select cameras and perform cuts between cameras that follow best practices in cinema, and (iii) we evaluate the benefits of FEPs by performing user experiments in which professional filmmakers and amateurs had to create cinematographic sequences. The evaluation suggests that users generally appreciate the idea of FEPs, and that it can effectively help novice and medium experienced users in crafting film sequences with little training. Hui-Yin Wu, Francesca Palù, Roberto Ranon, Marc Christie |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2016 | Trip Synopsis: 60km in 60secabstractAbstract Computerized route planning tools are widely used today by travelers all around the globe, while 3D terrain and urban models are becoming increasingly elaborate and abundant. This makes it feasible to generate a virtual 3D flyby along a planned route. Such a flyby may be useful, either as a preview of the trip, or as an after‐the‐fact visual summary. However, a naively generated preview is likely to contain many boring portions, while skipping too quickly over areas worthy of attention. In this paper, we introduce 3D trip synopsis: a continuous visual summary of a trip that attempts to maximize the total amount of visual interest seen by the camera. The main challenge is to generate a synopsis of a prescribed short duration, while ensuring a visually smooth camera motion. Using an application‐specific visual interest metric, we measure the visual interest at a set of viewpoints along an initial camera path, and maximize the amount of visual interest seen in the synopsis by varying the speed along the route. A new camera path is then computed using optimization to simultaneously satisfy requirements, such as smoothness, focus and distance to the route. The process is repeated until convergence. The main technical contribution of this work is a new camera control method, which iteratively adjusts the camera trajectory and determines all of the camera trajectory parameters, including the camera position, altitude, heading, and tilt. Our results demonstrate the effectiveness of our trip synopses, compared to a number of alternatives. Hui Huang 0004, Dani Lischinski, Hao (Richard) Zhang, Minglun Gong, Marc Christie, Daniel Cohen-Or |
Comput. Graph. Forum | 5 |
| 2016 | A Virtual Director Using Hidden Markov ModelsabstractAbstract Automatically computing a cinematographic consistent sequence of shots over a set of actions occurring in a 3D world is a complex task which requires not only the computation of appropriate shots (viewpoints) and appropriate transitions between shots (cuts), but the ability to encode and reproduce elements of cinematographic style. Models proposed in the literature, generally based on finite state machine or idiom‐based representations, provide limited functionalities to build sequences of shots. These approaches are not designed in mind to easily learn elements of cinematographic style, nor do they allow to perform significant variations in style over the same sequence of actions. In this paper, we propose a model for automated cinematography that can compute significant variations in terms of cinematographic style, with the ability to control the duration of shots and the possibility to add specific constraints to the desired sequence. The model is parametrized in a way that facilitates the application of learning techniques. By using a Hidden Markov Model representation of the editing process, we demonstrate the possibility of easily reproducing elements of style extracted from real movies. Results comparing our model with state‐of‐the‐art first‐order Markovian representations illustrate these features, and robustness of the learning technique is demonstrated through cross‐validation. Billal Merabti, Marc Christie, Kadi Bouatouch |
Comput. Graph. Forum | 2 |
| 2015 | Continuity Editing for 3D AnimationabstractWe describe an optimization-based approach for automatically creating well-edited movies from a 3D animation. While previous work has mostly focused on the problem of placing cameras to produce nice-looking views of the action, the problem of cutting and pasting shots from all available cameras has never been addressed extensively. In this paper, we review the main causes of editing errors in literature and propose an editing model relying on a minimization of such errors. We make a plausible semi-Markov assumption, resulting in a dynamic programming solution which is computationally efficient. We also show that our method can generate movies with different editing rhythms and validate the results through a user study. Combined with state-of-the-art cinematography, our approach therefore promises to significantly extend the expressiveness and naturalness of virtual movie-making. Quentin Galvane, Rémi Ronfard, Christophe Lino, Marc Christie |
AAAI | 4 |
| 2015 | Camera-on-rails: automated computation of constrained camera pathsabstractWhen creating real or computer graphics movies, the questions of how to layout elements on the screen, together with how to move the cameras in the scene are crucial to properly conveying the events composing a narrative. Though there is a range of techniques to automatically compute camera paths in virtual environments, none have seriously considered the problem of generating realistic camera motions even for simple scenes. Among possible cinematographic devices, real cinematographers often rely on camera rails to create smooth camera motions which viewers are familiar with. Following this practice, in this paper we propose a method for generating virtual camera rails and computing smooth camera motions on these rails. Our technique analyzes characters motion and user-defined framing properties to compute rough camera motions which are further refined using constrained-optimization techniques. Comparisons with recent techniques demonstrate the benefits of our approach and opens interesting perspectives in terms of creative support tools for animators and cinematographers. Quentin Galvane, Marc Christie, Christophe Lino, Rémi Ronfard |
MIG | 2 |
| 2015 | Crowd art: density and flow based crowd motion designabstractArtists, animation and game designers are in demand for solutions to easily populate large virtual environments with crowds that satisfy desired visual features. This paper presents a method to intuitively populate virtual environments by specifying two key features: localized density, being the amount of agents per unit of surface, and localized flow, being the direction in which agents move through a unit of surface. The technique we propose is also time-independant, meaning that whatever the time in the animation, the resulting crowd satisfies both features. To achieve this, our approach relies on the Crowd Patches model. After discretizing the environment into regular patches and creating a graph that links these patches, an iterative optimization process computes the local changes to apply on each patch (increasing/reducing the number of agents in each patch, updating the directions of agents in the patch) in order to satisfy overall density and flow constraints. A specific stage is then introduced after each iteration to avoid the creation of local loops by using a global pathfinding process. As a result, the method has the capacity of generating large realistic crowds in minutes that endlessly satisfy both user specified densities and flow directions, and is robust to contradictory inputs. At last, to ease the design the method is implemented in an artist-driven tool through a painting interface. Kevin Jordao, Panayiotis Charalambous, Marc Christie, Julien Pettré, Marie-Paule Cani |
MIG | 3 |
| 2015 | Intuitive and efficient camera control with the toric spaceabstractA large range of computer graphics applications such as data visualization or virtual movie production require users to position and move viewpoints in 3D scenes to effectively convey visual information or tell stories. The desired viewpoints and camera paths are required to satisfy a number of visual properties ( e.g. size, vantage angle, visibility, and on-screen position of targets). Yet, existing camera manipulation tools only provide limited interaction methods and automated techniques remain computationally expensive. In this work, we introduce the Toric space , a novel and compact representation for intuitive and efficient virtual camera control. We first show how visual properties are expressed in this Toric space and propose an efficient interval-based search technique for automated viewpoint computation. We then derive a novel screen-space manipulation technique that provides intuitive and real-time control of visual properties. Finally, we propose an effective viewpoint interpolation technique which ensures the continuity of visual properties along the generated paths. The proposed approach (i) performs better than existing automated viewpoint computation techniques in terms of speed and precision, (ii) provides a screen-space manipulation tool that is more efficient than classical manipulators and easier to use for beginners, and (iii) enables the creation of complex camera motions such as long takes in a very short time and in a controllable way. As a result, the approach should quickly find its place in a number of applications that require interactive or automated camera control such as 3D modelers, navigation tools or 3D games. Christophe Lino, Marc Christie |
ACM Trans. Graph. | 2 |
| 2014 | Interactive Design of Sustainable Cities with a Distributed Local Search Solver
Bruno Belin, Marc Christie, Charlotte Truchet |
CPAIOR | 2 |
| 2014 | Narrative-driven camera control for cinematic replay of computer gamesabstractThis paper presents a system that generates cinematic replays for dialogue-based 3D video games. The system exploits the narrative and geometric information present in these games and automatically computes camera framings and edits to build a coherent cinematic replay of the gaming session. We propose a novel importance-driven approach to cinematic replay. Rather than relying on actions performed by characters to drive the cinematography (as in idiom-based approaches), we rely on the importance of characters in the narrative. We first devise a mechanism to compute the varying importance of the characters. We then map importances of characters with different camera specifications, and propose a novel technique that (i) automatically computes camera positions satisfying given specifications, and (ii) provides smooth camera motions when transitioning between different specifications. We demonstrate the features of our system by implementing three camera behaviors (one for master shots, one for shots on the player character, and one for reverse shots). We present results obtained by interfacing our system with a full-fledged serious game (Nothing for Dinner) containing several hours of 3D animated content. Quentin Galvane, Rémi Ronfard, Marc Christie, Nicolas Szilas |
MIG | 3 |
| 2014 | Crowd sculpting: A space-time sculpting method for populating virtual environmentsabstractAbstract We introduce “Crowd Sculpting”: a method to interactively design populated environments by using intuitive deformation gestures to drive both the spatial coverage and the temporal sequencing of a crowd motion. Our approach assembles large environments from sets of spatial elements which contain inter‐connectible, periodic crowd animations. Such a “Crowd Patches” approach allows us to avoid expensive and difficult‐to‐control simulations. It also overcomes the limitations of motion editing, that would result into animations delimited in space and time. Our novel methods allows the user to control the crowd patches layout in ways inspired by elastic shape sculpting: the user creates and tunes the desired populated environment through stretching, bending, cutting and merging gestures, applied either in space or time. Our examples demonstrate that our method allows the space‐time editing of very large populations and results into endless animation, while offering real‐time, intuitive control and maintaining animation quality. Kevin Jordao, Julien Pettré, Marc Christie, Marie-Paule Cani |
Comput. Graph. Forum | 3 |
| 2013 | Steering Behaviors for Autonomous CamerasabstractThe automated computation of appropriate viewpoints in complex 3D scenes is a key problem in a number of computer graphics applications. In particular, crowd simulations create visually complex environments with many simultaneous events for which the computation of relevant viewpoints remains an open issue. In this paper, we propose a system which enables the conveyance of events occurring in complex crowd simulations. The system relies on Reynolds' model of steering behaviors to control and locally coordinate a collection of camera agents similar to a group of reporters. In our approach, camera agents are either in a scouting mode, searching for relevant events to convey, or in a tracking mode following one or more unfolding events. The key benefit, in addition to the simplicity of the steering rules, holds in the capacity of the system to adapt to the evolving complexity of crowd simulations by self-organizing the camera agents to track interesting events. Quentin Galvane, Marc Christie, Rémi Ronfard, Chen Kim Lim, Marie-Paule Cani |
MIG | 2 |
| 2012 | HapSeat: producing motion sensation with multiple force-feedback devices embedded in a seatabstractWe introduce a novel way of simulating sensations of motion which does not require an expensive and cumbersome motion platform. Multiple force-feedbacks are applied to the seated user's body to generate a sensation of motion experiencing passive navigation. A set of force-feedback devices such as mobile armrests or headrests are arranged around a seat so that they can apply forces to the user. We have dubbed this new approach HapSeat. A proof of concept has been designed which uses three low-cost force-feedback devices, and two control models have been implemented. Results from the first user study suggest that subjective sensations of motion are reliably generated using either model. Our results pave the way to a novel device to generate consumer motion effects based on our prototype. Fabien Danieau, Julien Fleureau, Philippe Guillotel, Nicolas Mollet, Anatole Lécuyer, Marc Christie |
VRST | 6 |
| 2011 | Computational Model of Film Editing for Interactive Storytelling
Christophe Lino, Mathieu Chollet, Marc Christie, Rémi Ronfard |
ICIDS | 3 |
| 2011 | The director's lens: an intelligent assistant for virtual cinematographyabstractWe present the Director's Lens, an intelligent interactive assistant for crafting virtual cinematography using a motion-tracked hand-held device that can be aimed like a real camera. The system employs an intelligent cinematography engine that can compute, at the request of the filmmaker, a set of suitable camera placements for starting a shot. These suggestions represent semantically and cinematically distinct choices for visualizing the current narrative. In computing suggestions, the system considers established cinema conventions of continuity and composition along with the filmmaker's previous selected suggestions, and also his or her manually crafted camera compositions, by a machine learning component that adapts shot editing preferences from user-created camera edits. The result is a novel workflow based on interactive collaboration of human creativity with automated intelligence that enables efficient exploration of a wide range of cinematographic possibilities, and rapid production of computer-generated animated movies. Christophe Lino, Marc Christie, Roberto Ranon, William H. Bares |
ACM Multimedia | 2 |
| 2011 | A smart assistant for shooting virtual cinematography with motion-tracked camerasabstractThis demonstration shows how an automated assistant encoded with knowledge of cinematography practice can offer suggested viewpoints to a filmmaker operating a hand-held motion-tracked virtual camera device. Our system, called Director's Lens, uses an intelligent cinematography engine to compute, at the request of the filmmaker, a set of suitable camera placements for starting a shot that represent semantically and cinematically distinct choices for visualizing the current narrative. Editing decisions and hand-held camera compositions made by the user in turn influence the system's suggestions for subsequent shots. The result is a novel virtual cinematography workflow that enhances the filmmaker's creative potential by enabling efficient exploration of a wide range of computer-suggested cinematographic possibilities. Christophe Lino, Marc Christie, Roberto Ranon, William H. Bares |
ACM Multimedia | 2 |
| 2008 | A Branch and Bound Algorithm for Numerical MAX-CSP
Jean-Marie Normand, Alexandre Goldsztejn, Marc Christie, Frédéric Benhamou |
CP | 3 |
| 2008 | A Tabu Search Method for Interval Constraints
Charlotte Truchet, Marc Christie, Jean-Marie Normand |
CPAIOR | 2 |
| 2008 | The IRIS Network of Excellence: Integrating Research in Interactive Storytelling
Marc Cavazza, Stéphane Donikian, Marc Christie, Ulrike Spierling, Nicolas Szilas, Peter Vorderer, Tilo Hartmann, Christoph Klimmt, Elisabeth André, Ronan Champagnat, Paolo Petta, Patrick Olivier |
ICIDS | 3 |
| 2008 | Camera Control in Computer GraphicsabstractAbstract Recent progress in modelling, animation and rendering means that rich, high fidelity virtual worlds are found in many interactive graphics applications. However, the viewer's experience of a 3D world is dependent on the nature of the virtual cinematography, in particular, the camera position, orientation and motion in relation to the elements of the scene and the action. Camera control encompasses viewpoint computation, motion planning and editing. We present a range of computer graphics applications and draw on insights from cinematographic practice in identifying their different requirements with regard to camera control. The nature of the camera control problem varies depending on these requirements, which range from augmented manual control (semi‐automatic) in interactive applications, to fully automated approaches. We review the full range of solution techniques from constraint‐based to optimization‐based approaches, and conclude with an examination of occlusion management and expressiveness in the context of declarative approaches to camera control. Marc Christie, Patrick Olivier, Jean-Marie Normand |
Comput. Graph. Forum | 1 |
| 2005 | A Semantic Space Partitioning Approach to Virtual Camera CompositionabstractPositioning a virtual camera in a 3D virtual environment is generally a non-intuitive task when using a 2D input device such as a mouse. The user has a mental representation of the result in terms of what he wants to see on the screen, and has to apply a mental inversion process to determine the location, orientation and field of view parameters of the camera. This task is generally achieved through a tedious and time-consuming process requiring a succession of "place the camera" and "check the result" operations. Current 3D modelers surprisingly lack integration of tools to assist the user in this task, despite the fact that cinema, in more than a hundred years, has provided a rich grammar that allows a director to unambiguously describe shots. Modelers are based on complex mathematical notions (spline curves, velocity graphs) more or less hidden by high-level manipulators. Manipulators allow positioning and animating the camera, but lack correlation with well established cinematographic notions relative to camera composition (object framing, distance shot specification, relative viewing angles, occlusions). In this paper we propose a new approach to virtual camera composition that fully relies on this grammar, relieves the user from low-level parameter manipulation and offers him classes of possible solutions w.r.t. current cinematographic notions. In related literature, numerous approaches have utilized cinematographic properties such as subject size and location within the frame to assist users in camera composition and camera planning. J. Blinn [Bli88] propose a vector algebra-based solution to pin two objects at given locations on the screen. Gleicher and Witkin [GW92] offer means to control 2D points directly on the display screen through differential manipulation techniques instead of controlling the camera parameters. Blinn's algebraic method has laid the groundwork for cinematographic idiom-based approaches, such as Christianson et al.'s compiler for the Declarative Camera Control Language (DCCL) [CAH*96] or He's et al.'s Virtual Cinematographer [HCS96]. These approaches however suffer from the point representation of the objects that doesn't capture the occlusion properties and thus avoids one of the main problems in camera planning, namely occlusion. Moreover the "point-like" representation cannot take into account the real geometry of the scene. Camera composition and camera planning can be viewed as constrained optimization problems in which the properties of the shot are expressed as numerical constraints on the camera variables (respectively the camera's path variables) and a broad range of solving procedures is available to compute solutions. The solvers differ in the way to manage over-constrained and under-constrained cases, in their complete or incomplete search capacities, in local minima management and possible optimization processes (generally finding the best solution w.r.t. an objective function). In their CAMDROID system for automated camera planning [DZ95], Drucker et al. use an existing numerical constraint solver package (CFSQP) and have to compile the shot specifications in the input script used by this library. Unfortunately, the solving process is sensitive to the initial configuration and is subject to local minima failures. Bares et al. use a partial constraint satisfaction system named CONSTRAINTCAM [BGL98] in order to provide alternate solutions when constraints cannot be completely satisfied. This solution is based on a limited subset of cinematographic properties (viewing angle, viewing distance and occlusion avoidance), which limits the procedure to small problems. P. Olivier et al.[OHPL99] propose a rich set of properties for composition purposes. The authors have developed the CAMPLAN system [HO00] that numerically solves the optimization problem via a metaheuristic search (genetic algorithms) method. The main shortcomings of this purely optimization-based technique is that the CAMPLAN genetic algorithm produces solutions in widely varying amounts of time and is also subject to the initial population of solutions. Gooch et al.[GRMS01] also use optimization procedures in order to produce images fitting artistic composition criterion like rules of thirds and fifths and the notion of canonical viewpoint (viewpoint leading to better identification of an object). The CSP (Constraint Satisfaction Problem) framework has proven to succeed in some camera composition and motion planning approaches. In [BMBT00] Bares et al. propose a heuristic-based complete search algorithm. The process is applied inside promising 3D areas computed through simple geometric intersections. Although efficient, the approach avoids problems related to multiple solutions. Unlike Bares et al. who utilize partial constraint satisfaction through cost functions, Jardillier & Languénou use pure interval methods in The Virtual Cameraman[JL98] to compute camera paths which yield sequences of images fulfilling temporally indexed image properties. This idea has been improved by Christie et al. in [CLG02]. Unfortunately, the benefits of interval-based techniques that guarantee the fulfillment of the properties during the whole sequence are counterbalanced by the computational effort required, and the absence of any mechanism for constraint relaxation. However, whenever the method fails, the user has a guarantee that there are no solutions to his problem (due to completeness of interval-based approaches). Some approaches combine constraints and optimization. One of these was presented by J. Pickering [Pic02] and is closely related to our work. In this solving method, the constraints are used to create feasible regions of space that will serve as bounds for an optimization procedure. The search space is subdivided by a "shadow-volumes" algorithm based on the properties of the image, and the feasible regions are then discretized and stored in an octree structure. Each node of the octree is then used as a starting point for a genetic algorithm that tries to find a solution to the problem. The main shortcomings of this approach lay in the computational effort required to create the octree, and the fact that multiple distinct solutions are ignored. Most of the methods we mention rely upon optimization processes to compute satisfactory camera placements and therefore lead to a unique solution closely related to the objective function. However, the description of a cinematic shot can possibly yield different visual solutions. Therefore, in computing the set of semantically distinct solutions w.r.t. cinematographic properties, one provides the user meaningful results. Figure 1 presents a top view of a simple scene containing three objects A, B and C. Whenever the user describes a shot in which he constrains A and B respectively to lay on the left and on the right of the screen, it clearly yields three possible classes of camera configurations: area (1) object C is on the left of A and B on the screen, (2) object C is between A and B, and (3) object C is on the right of A and B. Moreover, when considering possible occlusions, two classes can be added (4) A occludes C and (5) B occludes C. In such cases, classical optimization and incomplete CSP-based approaches fail in that a unique solution [DZ95,OHPL99,Pic02], or a reduced subset [JL98,CLG02] of solutions is proposed, whereas all classes of solutions should be equally considered. These approaches actually lose the semantics of the problem while relying upon pure numerical approaches. Once a solution is computed, no further information on its characteristics or differences with other possible solutions is provided. Certainly, any classification process can be provided afterhand, but requires an important computational effort. Possible distinct areas for viewing a couple of objects A and B (resp. on the left and right of the screen) w.r.t. to a third object C. In this paper, we propose to integrate a semantic dimension in the solving and interaction processes to assist the user in his camera placement tasks. We follow a threefold declarative approach: describe the desired solution with a set of cinematographic-based properties, compute distinct classes of solutions satisfying the description with related cinematographic properties, explore and interact with the classes of possible solutions. In the description phase, as in previous approaches [OHPL99,HO00,JL98,DZ95], a high-level grammar is offered including composition properties (framing objects on screen surface, relative object orientation and size) and shot properties (close shot, establishing shot, low and high angle). The geometry of the scene (locations and orientation of objects) is considered as an input provided by the user. In the computational phase, two processes are combined. The first process partitions the search space according to cinematographic properties (e.g. area such that A occludes B on the screen) and builds the intersection of the space partitions. The second process computes a nice representative of each possible class of solutions via a continuous domain implementation of a local search metaheuristic algorithm. Finally, in a third phase, the user navigates in the possible solution sets and interacts with the semantic information provided in each area. This paper concentrates on the first two phases and offers solid foundations for high-level interactions with the user. This paper is organized as follows: Section 2 introduces our semantic space partitioning approach to virtual camera composition, Section 3 details the numerical solving process. The exploitation of the semantic volumes is presented in Section 4 and relevant results are then presented in Section 5. Finally Section 6 discusses future research directions and concludes. Our approach to virtual camera compostion (VCC) is based on the primary idea of Binary Space Partition (BSP) and can be considered as an extension of visual aspects[KvD79] and closely related works such as viewpoint space partitioning[PD90] in the field of object recognition. The idea behind visual aspects is to gather all the viewpoints of a single polyhedron that share similar topological characteristics on the image. A change of appearance of the polyhedron with changing viewpoint, gives rise to boundaries in the search space. Computing all the boundaries enables the construction of regions of constant aspect, namely viewpoint space partitions. In this paper, we propose an extension of viewpoint space partitions to multiple objects and replace the topological characteristics of a polyhedron by cinematographic properties such as occlusions, relative viewing angles, distance shots and relative object locations. We introduce the notion of semantic volume as a volume of possible camera locations that give rise to qualitatively equivalent shots w.r.t. to cinematographic properties, i.e. semantically equivalent shots. Each volume is characterized by a set of semantic tags issued from film grammar [Ari76] and each tag is associated to a satisfied property in the volume. Tags are either related to a single object such as viewing angle, or to a couple of objects such as occlusion and relative image location. Therefore, the entire space of possible camera locations is thoroughly partitioned for each object, and each couple of objects. We then derive from the user's description the subset of volumes to be considered and intersect them. This process leads to a set of non-connected regions of which each represents a different class of solutions in terms of visual aspect. The computation of semantic volumes can be formalized as follows. We define the geometric filtering operator Gƒ that inputs a property p and provides a semantic volume sv, which is defined by sv=〈S, V〉 where S is a conjunction of semantic tags (e.g. LeftOf(A) ΛMediumShotOn(B) ΛOccludes(A,B) Λ…) and V is a subset of possible camera locations in ℜ3. Every camera location inside sv possibly satisfies the property p, whereas every camera location outside sv certainly violates p. Possible camera orientations are to be further computed by the numerical process (see Section 3). The filtering operator is complete in that it does not lose any correct camera locations. The operator Gƒ computes (possibly) non-connected volumes by pruning the space of most of the inconsistent camera locations w.r.t. p and associates a semantic tag to the resulting volumes. For example, the user description "A occludes B" (see Fig. 5) leads to a cone-shaped volume V coupled with the semantic tag Occlusion (B,A). Occlusion cones computation (both partial and total occlusions). The following subsections present how the filtering operator Gf provides the semantic volumes related to each property. Below, we consider that the camera's roll degree of freedom (i.e. around the look-at vector) is restricted to interval , and the tilt angle is confined in (no upside-down cameras). Camera compositions seldom break these conventions. The projection property is based on the notion of scale shots in cinematography (cf.Fig 2). It allows the artist to specify a viewing shot for an object. There are basically six different kinds of shots : the Extreme Close-Up, the Close-Up, the Medium Close-Up, the Medium Long Shot (or Plan Américain), the Long Shot, the Extreme Long Shot. Related semantic tags reflect all six kinds of shots (ExtremeCloseUp(Object) to ExtremeLongShot(Object)). The six distances between the camera and a character according to Arijon [Ari76]. The underlying semantic volume is computed given the position of an object and a cinematographic scale shot specified by the user. One can deduce an optimal size corresponding to each scale shot presented in Figure 2. The object's bounding sphere and desired area in the frame are used to determine the range of camera distances. In order to add some flexibility to the solving system, we compute an interval range related to this optimal value by subtracting and adding an epsilon to it (see [BMBT00]). The minimum and maximum bounds of the interval correspond to two distances defining the inner and outer radiuses of an hollow sphere that includes the set of consistent positions for the camera (cf.Fig. 3). Distance is trivially computed by the following equation: Semantic volumes leading to characteristics shots. The orientation property lets the virtual cinematographer specify the viewing angle required to shoot an object or a character. A common set of 8 viewing angles is offered (e.g. relative to object A, there are IsLeftProfileOf(A), IsRightProfileOf(A), IsInFrontOf(A), IsInBackOf(A), IsThreeQuaterFrontLeft(A) up to IsThreeQuaterBackRight(A)) and each can be composed with high and low relative angles IsHighAngle(A) and IsLowAngle(A). Computing orientation semantic volumes consists in building a prism-shaped volume of possible camera locations, w.r.t. the vector to consider (front, back, left, …). Once again, a variation with the optimal orientation is accepted in order to avoid being too restrictive (cf.Fig. 4). Four common relative viewing angles and related semantic tags. The occlusion property gives the user the opportunity to specify some visibility constraints between two objects of the scene. The cinematographer can characterize a total occlusion of an object by another, a partial occlusion or an absence of occlusion between two objects. A partial occlusion occurs when a part of an object's projection overlaps some part of the second object's projection. The semantic volumes induced by an occlusion property are computed given characteristic cones defined with respect to the positions of the two objects involved in the occlusionproperty [DDP02]. The inside bounds of the cones define the volumes of partial occlusion. The outside bounds of the cones define the volumes where no possible occlusion can occur (cf.Fig. 5). The framing property allows the virtual cinematographer to constrain an object in a given frame inside, partially inside, or outside the screen space. This expressive property defines the relative locations and sizes of objects on the screen. It constrains altogether the distance shot and the camera locations and orientations. Moreover, total or partial occlusions can be derived from overlapping frames, and conversely non-occlusions can be derived from the absence of overlap (cf.Fig. 6and 8). A non-overlapping description constraining three objects. An overlapping description constraining three objects with an overlap. The computation and characterization of the semantic volumes related to this property depend on the number of framing properties defined by the user and their relative location on the screen. We two (1) whenever a single frame is the framing property will like a projection the size of the frame and the size of the object in the scene lead to the computation of a volume that limits the all possible camera locations around the object are possible w.r.t. this (2) the user has specified two or more a is to be for every couple of objects in order to The of all the distinct frame is not provided we propose a two leading to different semantic volumes considering overlapping and non-overlapping 8 and that the semantic volumes computed at the of the framing will the between the on the screen (e.g. object A is left of but not the locations of the objects in the the numerical process will compute camera orientations to locations in the Figure 6 a user input related to a non-overlapping Figure provides a 2D of the scene with locations of objects A, B and C. For the representation is two but all the computation occurs in The first semantic volume is as 1 A and and with possible camera location in this area can lead to a shot in which the object A on the left of B. as 1 and 2 not overlap on the screen, no occlusion should occur between A and B. 2 and 3 are by computing the occlusion cones [DDP02]. tags are where B A and when A B. the area of Figure for the possible camera locations satisfying the user's description 2D of the semantic volumes related to couple , when respectively on the left and right of the screen. This process is for every couple of objects in the screen, Figure 8 an overlapping the overlapping is partial and areas and are not considered as possible camera locations and are with or the area between A and B is not considered as possible 3 does not provide any overlapping configuration and is with in the areas containing possible camera locations are at the left and right of Figure distinct volumes are the camera is in the right solution B will than A in the shot, whereas it is to the left, A will than B. 2D of semantic volumes related to overlapping on objects A and B the further semantic partitioning and characterization is offered by considering and The of the whole set of distinct frame follow the process. positioning properties allow to specify some relative positions between objects on the screen, one can that one wants to see the object A on the right of B, B, These relative placements describe the of objects being to within a visual composition as a conjunction of properties is by a intersection of the semantic volumes. a set of properties provided by the user results in the following computation : The intersection process lead to an a unique volume or to a set of non-connected volumes. In order to the intersection of 3D we propose to rely on of the semantic volumes than pure geometric Each is the of a field the function. For each point of the the of the are A 3D volume is when a value such that provides with the following (1) of volumes by the (2) in that we avoid or (3) to a point in the via a and simple Our implementation relies upon the that provides and means to create by defining for each object and with the bounding the whole 3D scene. For example, the possible camera locations the scale shots distances between the camera and an object (cf.Fig. are via the volumes the occlusion property are defined as cones the objects (see Fig. thus the regions of total or partial The to the use of in the computational cost of our approach requires a unique at the in order to determine the number of non-connected The cost of the is by the number of and for our the computation of the non-connected with The of the space partitioning approach consists in a semantic volume V containing possible camera locations. the numerical computes a nice representative of each volume in This consists in in V a consistent camera configuration orientation and that satisfies the framing properties and that each property corresponding to a semantic tag of S a cost function). The problem therefore to determine a of variables such that : where for the cost associated to property and where is the framing property given by the user. optimization and constrained optimization techniques a objective (e.g. In order to manage algebraic constraints such as and algebraic constraints such as we rely on The underlying of the framework can be in 1 presents our The algorithm relies on the of the search space a semantic starting from an initial and the around the current The a set of in V within a around the current It introduces the notions of and The allows on promising regions of the search space by the size the being in local no has been for a while a of the initial is to allow a of the search space. The procedure is by the maximum number of and and the number of at each of the function. each the best in terms of constraint satisfaction and cost the new current the user's description not lead to a solution framing is to respect as objects have been in the 3D we propose to integrate the satisfaction of the framing constraints in the cost function. Therefore, is given by : The cost to the respect of a property considering the orientation property of for example, the best are vector and such that the camera's orientation is to This local search technique be used to compute solutions of camera composition problems to Olivier et al.'s However, the process the solver in promising regions and the of the associated to each property. The constraint is by a simple of the related to a semantic the configuration in V and conversely is outside The main of this paper is to offer a semantic for and with the volumes. We two possible on the computed semantic volumes and on the whole 3D scene the computed volumes. For a given each computed volume provided by the geometric solver some related to the satisfaction of the properties. the characterization of each distinct volume can be semantically according to properties the user has not For example, a semantic volume sv is characterized by the relative location some further characterization of sv can be computed by considering the orientation properties related to A and B (e.g. a the user can two semantically volumes and the description and for the differences between them. This to compute all tags in and in that not each object and couple of objects in the scene their semantic it is to on a computed volume sv any possible object or property. The to computing a new geometric intersection and the number of computed by For example, in 2 of Section one can there a possible camera location such that A, B and C can be viewed from the relative angle : where sv is the computed semantic volume and Gf the geometric that computes the semantic volume related to the property is the is all possible volumes the a computed volume w.r.t. a property a similar For example, the of a volume sv into all the possible shot distances relative to an object A can be expressed as : The set of properties in Section 2 is provided via a as an extension to the It is for the and the of the solutions given to the user. Moreover it the into existing by the of and 2 the related to most properties. The framing property as parameters the of the left and top right of the frame containing the object. less than or more than 1 frame objects of the screen. The orientation property requires an object and a viewing angle presented in Section The projection property an object and a shot class 2). Occlusion properties associates with and the occlusion or our the classical shot two and then a framing shot with possible classes of solutions. The classical shot is when a between two or more and consists in the camera behind one while framing the Figure the related semantic volume and a result is presented in Figure The shot is given by the following declarative script w.r.t. and view of the search space computed by the shot and of the volume A result of the The geometry of the scene is composed of objects (see Fig. The user objects A, B and C respectively in the left, and right of the screen, and constrains and to to the screen any occlusion. Some results are presented in Figure three shots the users description and different classes of solutions. view of the shot and related to possible camera locations with further semantic information related to orientations of A , B and C shots associated to the framing 3 presents the time during the and numerical in computation of one representative of each semantic volume. is directly related to the of the are and 1 and 2 respectively and is the time in the local search with and provide a Although the total time to compute all important time representative is around which is for interaction purposes. The semantic approach offers the following the cinematographic properties provide semantic volumes containing the possible solutions of the problem through a geometric process that areas of the The computation of the boundaries relies on and avoids volume Whenever the intersection process leads to an there is a guarantee of in the user's The numerical process offers a representative of each volume at low computational The use of as object boundaries in the occlusion computation can be present the of being for bounding objects that are one or for This lead to regions of the search space that possibly lead to correct shots during occlusions and results can be through but to integrate in our can be by computing volumes provided an representation of such volumes is The computational cost of our approach is related to the number of objects in the scene and to the user's Most time is in the computation of the number of non-connected volumes which is related to the of the function. We to this process by computing the intersection of of each semantic volume and then the on this Finally, the approach to is our main objective being to characterize possible camera positions for a of time the computation of camera A extension is volumes. the extension is not and the main in with the time criterion in each volume. in this paper we have presented an approach to virtual camera composition that classes of distinct provides means to characterize and computes the notion of visual we the notion of Semantic as a set of possible camera locations that share a set of cinematographic results the of our approach and in and to virtual camera Marc Christie, Jean-Marie Normand |
Comput. Graph. Forum | 1 |
| 2004 | Interval constraint solving for camera control and motion planningabstractMany problems in robust control and motion planning can be reduced to either finding a sound approximation of the solution space determined by a set of nonlinear inequalities, or to the "guaranteed tuning problem" as defined by Jaulin and Walter, which amounts to finding a value for some tuning parameter such that a set of inequalities be verified for all the possible values of some perturbation vector. A classical approach to solving these problems, which satisfies the strong soundness requirement, involves some quantifier elimination procedure such as Collins' Cylindrical Algebraic Decomposition symbolic method. Sound numerical methods using interval arithmetic and local consistency enforcement to prune the search space are presented in this article as much faster alternatives for both soundly solving systems of nonlinear inequalities, and addressing the guaranteed tuning problem whenever the perturbation vector has dimension 1. The use of these methods in camera control is investigated, and experiments with the prototype of a declarative modeler to express camera motion using a cinematic language are reported and commented upon. Frédéric Benhamou, Frédéric Goualard, Éric Languénou, Marc Christie |
ACM Trans. Comput. Log. | 4 |
| 2002 | Modeling Camera Control with Constrained Hypertubes
Marc Christie, Éric Languénou, Laurent Granvilliers |
CP | 1 |