Xi Wang 0024

dblp:08/5760-24 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0001-6586-1926ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2025 AKiRa: Augmentation Kit on Rays for Optical Video Generation
abstract
Recent advances in text-conditioned video diffusion have greatly improved video quality. However, these methods offer limited or sometimes no control to users on camera aspects, including dynamic camera motion, zoom, distorted lens and focus shifts. These motion and optical aspects are crucial for adding controllability and cinematic elements to generation frameworks, ultimately resulting in visual content that draws focus, enhances mood, and guides emotions according to filmmakers’ controls. In this paper, we aim to close the gap between controllable video generation and camera optics. To achieve this, we propose AKiRa (Augmentation Kit on Rays), a novel augmentation framework that builds and trains a camera adapter with a complex camera model over an existing video generation backbone. It enables fine-tuned control over camera motion as well as complex optical parameters (focal length, distortion, aperture) to achieve cinematic effects such as zoom, fish-eye effect, and bokeh. Extensive experiments demonstrate AKiRa’s effectiveness in combining and composing camera optics while outperforming all state-of-the-art methods. This work sets a new landmark in controlled and optically enhanced video generation, paving the way for future optical video generation methods.
Xi Wang 0024, Robin Courant, Marc Christie, Vicky Kalogeiton
CVPR1
2025 Di[M]O: Distilling Masked Diffusion Models Into One-Step Generator
Yuanzhi Zhu 0001, Xi Wang 0024, Stéphane Lathuilière, Vicky Kalogeiton
ICCV2
2025 LEAD: Latent Realignment for Human Motion Diffusion
abstract
Abstract Our goal is to generate realistic human motion from natural language. Modern methods often face a trade‐off between model expressiveness and text‐to‐motion (T2M) alignment. Some align text and motion latent spaces but sacrifice expressiveness; others rely on diffusion models producing impressive motions but lacking semantic meaning in their latent space. This may compromise realism, diversity and applicability. Here, we address this by combining latent diffusion with a realignment mechanism, producing a novel, semantically structured space that encodes the semantics of language. Leveraging this capability, we introduce the task of textual motion inversion to capture novel motion concepts from a few examples. For motion synthesis, we evaluate LEAD on HumanML3D and KIT‐ML and show comparable performance to the state‐of‐the‐art in terms of realism, diversity and text‐motion consistency. Our qualitative analysis and user study reveal that our synthesised motions are sharper, more human‐like and comply better with the text compared to modern methods. For motion textual inversion (MTI), our method demonstrates improvements in capturing out‐of‐distribution characteristics in comparison to traditional VAEs.
Nefeli Andreou, Xi Wang 0024, Victoria Fernández Abrevaya, Marie-Paule Cani, Yiorgos Chrysanthou, Vicky Kalogeiton
Comput. Graph. Forum2
2024 E.T. the Exceptional Trajectories: Text-to-Camera-Trajectory Generation with Character Awareness
Robin Courant, Nicolas Dufour, Xi Wang 0024, Marc Christie, Vicky Kalogeiton
ECCV (4)3
2024 Cinematographic Camera Diffusion Model
abstract
Abstract Designing effective camera trajectories in virtual 3D environments is a challenging task even for experienced animators. Despite an elaborate film grammar, forged through years of experience, that enables the specification of camera motions through cinematographic properties (framing, shots sizes, angles, motions), there are endless possibilities in deciding how to place and move cameras with characters. Dealing with these possibilities is part of the complexity of the problem. While numerous techniques have been proposed in the literature (optimization‐based solving, encoding of empirical rules, learning from real examples,…), the results either lack variety or ease of control. In this paper, we propose a cinematographic camera diffusion model using a transformer‐based architecture to handle temporality and exploit the stochasticity of diffusion models to generate diverse and qualitative trajectories conditioned by high‐level textual descriptions. We extend the work by integrating keyframing constraints and the ability to blend naturally between motions using latent interpolation, in a way to augment the degree of control of the designers. We demonstrate the strengths of this text‐to‐camera motion approach through qualitative and quantitative experiments and gather feedback from professional artists. The code and data are available at https://github.com/jianghd1996/Camera-control .
Hongda Jiang, Xi Wang 0024, Marc Christie, Libin Liu 0002, Baoquan Chen
Comput. Graph. Forum2
2023 Real-time Computational Cinematographic Editing for Broadcasting of Volumetric-captured events: an Application to Ultimate Fighting
abstract
The capacity to capture and broadcast sports events with close to real-time volumetric reconstruction techniques opens exciting perspectives in how audiences can consume and interact with these contents. In this work, we propose the design of an real-time cinematography system that is capable of generating qualitative framing and editing of volumetric-captured content, by mimicking real broadcast footage. To illustrate our approach, we focus on the specific problem of cinematography for ring-based events such as Ultimate Fighting Championships (UFC). We start by extracting statistical features from hours of real footage to understand the specific framing and cutting behaviors of real broadcast directors. We then exploit these features in a real-time editing system, not only to replicate the behaviors in existing broadcasting, but also to generalize to novel camera layouts. We demonstrate our approach on the volumetric reconstruction of UFC fights, compare our results with baseline methods, and report different qualitative and quantitative evaluations.
Francois Bourel, Xi Wang 0024, Ervin Teng, Valerio Ortenzi, Adam Myhill, Marc Christie
MIG2
2021 Camera keyframing with style and control
abstract
We present a novel technique that enables 3D artists to synthesize camera motions in virtual environments following a camera style , while enforcing user-designed camera keyframes as constraints along the sequence. To solve this constrained motion in-betweening problem, we design and train a camera motion generator from a collection of temporal cinematic features (camera and actor motions) using a conditioning on target keyframes. We further condition the generator with a style code to control how to perform the interpolation between the keyframes. Style codes are generated by training a second network that encodes different camera behaviors in a compact latent space, the camera style space. Camera behaviors are defined as temporal correlations between actor features and camera motions and can be extracted from real or synthetic film clips. We further extend the system by incorporating a fine control of camera speed and direction via a hidden state mapping technique. We evaluate our method on two aspects: i) the capacity to synthesize style-aware camera trajectories with user defined keyframes; and ii) the capacity to ensure that in-between motions still comply with the reference camera style while satisfying the keyframe constraints. As a result, our system is the first style-aware keyframe in-betweening technique for camera control that balances style-driven automation with precise and interactive control of keyframes.
Hongda Jiang, Marc Christie, Xi Wang 0024, Libin Liu 0002, Bin Wang 0021, Baoquan Chen
ACM Trans. Graph.3