Robin Courant

dblp:339/2624 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0002-5329-4009ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Visual content generation and editing · 54% Computer animation and physical simulation · 17% Rendering · 15%
Artificial intelligence
3 papers
Vision and language · 37% Video understanding and tracking · 37% Generative modeling · 26%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing › video generation › controllable video generation
camera-controlled video generation
0.912025
AKiRa: Augmentation Kit on Rays for Optical Video Generation · CVPR 2025
Visual content generation and editing › video generation
controllable video generation
0.912025
AKiRa: Augmentation Kit on Rays for Optical Video Generation · CVPR 2025
Computer animation and physical simulation › virtual cinematography
camera trajectory generation
0.812024
E.T. the Exceptional Trajectories: Text-to-Camera-Trajectory Generation with Character Awareness · ECCV (4) 2024
Visual content generation and editing
camera control
0.712023
JAWS: Just A Wild Shot for Cinematic Transfer in Neural Radiance Fields · CVPR 2023
Computational photography and imaging
camera parameter optimization
0.712023
JAWS: Just A Wild Shot for Cinematic Transfer in Neural Radiance Fields · CVPR 2023
Rendering
neural radiance fields
0.712023
JAWS: Just A Wild Shot for Cinematic Transfer in Neural Radiance Fields · CVPR 2023
Machine learning › Generative modeling
diffusion model
0.312025
AKiRa: Augmentation Kit on Rays for Optical Video Generation · CVPR 2025
Machine learning › Generative modeling › diffusion model
video diffusion model
0.312025
AKiRa: Augmentation Kit on Rays for Optical Video Generation · CVPR 2025

Methods — techniques the papers use, named apart from their topics

diffusion model · 1.7camera adapter · 1.7speech-to-text · 0.8self-attention · 0.8large language model · 0.8cross-attention · 0.8implicit neural representation · 0.7differentiable rendering · 0.7backpropagation · 0.7
YearPublicationVenuePosition
2025 AKiRa: Augmentation Kit on Rays for Optical Video Generation
abstract
Recent advances in text-conditioned video diffusion have greatly improved video quality. However, these methods offer limited or sometimes no control to users on camera aspects, including dynamic camera motion, zoom, distorted lens and focus shifts. These motion and optical aspects are crucial for adding controllability and cinematic elements to generation frameworks, ultimately resulting in visual content that draws focus, enhances mood, and guides emotions according to filmmakers’ controls. In this paper, we aim to close the gap between controllable video generation and camera optics. To achieve this, we propose AKiRa (Augmentation Kit on Rays), a novel augmentation framework that builds and trains a camera adapter with a complex camera model over an existing video generation backbone. It enables fine-tuned control over camera motion as well as complex optical parameters (focal length, distortion, aperture) to achieve cinematic effects such as zoom, fish-eye effect, and bokeh. Extensive experiments demonstrate AKiRa’s effectiveness in combining and composing camera optics while outperforming all state-of-the-art methods. This work sets a new landmark in controlled and optically enhanced video generation, paving the way for future optical video generation methods.
Xi Wang 0024, Robin Courant, Marc Christie, Vicky Kalogeiton
CVPR2
2024 E.T. the Exceptional Trajectories: Text-to-Camera-Trajectory Generation with Character Awareness
Robin Courant, Nicolas Dufour, Xi Wang 0024, Marc Christie, Vicky Kalogeiton
ECCV (4)1
2024 FunnyNet-W: Multimodal Learning of Funny Moments in Videos in the Wild
abstract
Abstract Automatically understanding funny moments (i.e., the moments that make people laugh) when watching comedy is challenging, as they relate to various features, such as body language, dialogues and culture. In this paper, we propose FunnyNet-W, a model that relies on cross- and self-attention for visual, audio and text data to predict funny moments in videos. Unlike most methods that rely on ground truth data in the form of subtitles, in this work we exploit modalities that come naturally with videos: (a) video frames as they contain visual information indispensable for scene understanding, (b) audio as it contains higher-level cues associated with funny moments, such as intonation, pitch and pauses and (c) text automatically extracted with a speech-to-text model as it can provide rich information when processed by a Large Language Model. To acquire labels for training, we propose an unsupervised approach that spots and labels funny audio moments. We provide experiments on five datasets: the sitcoms TBBT, MHD, MUStARD, Friends, and the TED talk UR-Funny. Extensive experiments and analysis show that FunnyNet-W successfully exploits visual, auditory and textual cues to identify funny moments, while our findings reveal FunnyNet-W’s ability to predict funny moments in the wild. FunnyNet-W sets the new state of the art for funny moment detection with multimodal cues on all datasets with and without using ground truth information.
Robin Courant, Vicky Kalogeiton
Int. J. Comput. Vis.2
2023 JAWS: Just A Wild Shot for Cinematic Transfer in Neural Radiance Fields
abstract
This paper presents JAWS, an optimization-driven approach that achieves the robust transfer of visual cinematic features from a reference in-the-wild video clip to a newly generated clip. To this end, we rely on an implicit-neural-representation (INR) in a way to compute a clip that shares the same cinematic features as the reference clip. We propose a general formulation of a camera optimization problem in an INR that computes extrinsic and intrinsic camera parameters as well as timing. By leveraging the differentiability of neural representations, we can back-propagate our designed cinematic losses measured on proxy estimators through a NeRF network to the proposed cinematic parameters directly. We also introduce specific enhancements such as guidance maps to improve the overall quality and efficiency. Results display the capacity of our system to replicate well known camera sequences from movies, adapting the framing, camera parameters and timing of the generated video clip to maximize the similarity with the reference clip.
Robin Courant, Jinglei Shi, Éric Marchand, Marc Christie
CVPR2
2022 FunnyNet: Audiovisual Learning of Funny Moments in Videos
Robin Courant, Vicky Kalogeiton
ACCV (4)2