VLDB 2026 Research / reviewers in the wild / expert
Daniel Holden
dblp:165/9944
· DBLP profile ↗
17ranked-venue papers
5as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Control Operators for Interactive Character AnimationabstractNeural-network-based character controllers are increasingly common and capable. However, the integration of desired control inputs such as joystick movement, motion paths, and objects in the environment, remains challenging. This is because these inputs often require custom feature engineering, specific neural network architectures, and training procedures. This renders these methods largely inaccessible to non-technical designers. To address this challenge, we introduce Control Operators , a powerful and flexible framework for specifying the control mechanisms of interactive character controllers. By breaking down the control problem into a set of simple operators, each with a semantic meaning for designers, and a corresponding neural network structure, we allow non-technical users to design control mechanisms in a way that is intuitive and can be composed together to train models that have multiple skills and control modes. We demonstrate their potential with two current state-of-the-art interactive character controllers - a Flow-Matching-based auto-regressive model, and a variation of Learned Motion Matching. We validate the approach via a user study wherein industry practitioners with varying degrees of ML and technical expertise explore the use of our system. Ruiyu Gou, Michiel van de Panne, Daniel Holden |
ACM Trans. Graph. | 3 |
| 2024 | Humanlike Behavior in a Third-Person Shooter with Imitation LearningabstractWe tackle the problem of generating humanlike bot behavior by learning from human demonstrations. We developed a controlled gym environment to collect data on a subset of human behavior-namely aiming and target acquisition in single opponent settings. We introduce an identity-conditioned causal transformer to produce humanlike behavior of a controllable quality on a per-frame basis that captures the differences in skill and style between conditioned players. Alexander R. Farhang, Brendan Mulcahy, Daniel Holden, Iain Matthews, Yisong Yue |
CoG | 3 |
| 2024 | HUMOS: Human Motion Model Conditioned on Body Shape
Shashank Tripathi, Omid Taheri, Christoph Lassner, Michael J. Black, Daniel Holden, Carsten Stoll |
ECCV (16) | 5 |
| 2023 | ZeroEGGS: Zero-shot Example-based Gesture Generation from SpeechabstractAbstract We present ZeroEGGS, a neural network framework for speech‐driven gesture generation with zero‐shot style control by example. This means style can be controlled via only a short example motion clip, even for motion styles unseen during training. Our model uses a Variational framework to learn a style embedding, making it easy to modify style through latent space manipulation or blending and scaling of style embeddings. The probabilistic nature of our framework further enables the generation of a variety of outputs given the input, addressing the stochastic nature of gesture motion. In a series of experiments, we first demonstrate the flexibility and generalizability of our model to new speakers and styles. In a user study, we then show that our model outperforms previous state‐of‐the‐art techniques in naturalness of motion, appropriateness for speech, and style portrayal. Finally, we release a high‐quality dataset of full‐body gesture motion including fingers, with speech, spanning across 19 different styles. Our code and data are publicly available at https://github.com/ubisoft/ubisoft‐laforge‐ZeroEGGS . Saeed Ghorbani, Ylva Ferstl, Daniel Holden, Nikolaus F. Troje, Marc-André Carbonneau |
Comput. Graph. Forum | 3 |
| 2021 | Artist guided generation of video game production quality face texturesabstractWe develop a high resolution face texture generation system which uses artist provided appearance controls as the conditions for a generative network. Artists are able to control various elements in the generated textures, such as the skin, eye, lip, and hair color. This is made possible by reparameterizing our dataset to the same UV mapping, allowing us to utilize image-to-image translation networks. Although our dataset is limited in size, only 126 samples in total, our system is still able to generate realistic face textures which strongly adhere to the input appearance attribute conditions because of our training augmentation methods. Once our system has generated the face texture, it is ready to be used in a modern game production environment. Thanks to our novel SuperResolution and material property recovery methods, our generated face textures are 4K resolution and have the associated material property maps required for raytraced rendering. Christian Murphy, Sudhir P. Mudur, Daniel Holden, Marc-André Carbonneau, Donya Ghafourzadeh, Andre Beauchamp |
Comput. Graph. | 3 |
| 2021 | SuperTrack: motion tracking for physically simulated characters using supervised learningabstractIn this paper we show how the task of motion tracking for physically simulated characters can be solved using supervised learning and optimizing a policy directly via back-propagation. To achieve this we make use of a world model trained to approximate a specific subset of the environment's transition function, effectively acting as a differentiable physics simulator through which the policy can be optimized to minimize the tracking error. Compared to popular model-free methods of physically simulated character control which primarily make use of Proximal Policy Optimization (PPO) we find direct optimization of the policy via our approach consistently achieves a higher quality of control in a shorter training time, with a reduced sensitivity to the rate of experience gathering, dataset size, and distribution. Levi Fussell, Kevin Bergamin, Daniel Holden |
ACM Trans. Graph. | 3 |
| 2020 | Appearance Controlled Face Texture Generation for Video Game CharactersabstractManually creating realistic, digital human heads is a difficult and time-consuming task for artists. While 3D scanners and photogrammetry allow for quick and automatic reconstruction of heads, finding an actor who fits specific character appearance descriptions can be difficult. Moreover, modern open-world videogames feature several thousands of characters that cannot realistically all be cast and scanned. Therefore, researchers are investigating generative models to create heads fitting a specific character appearance description. While current methods are able to generate believable head shapes quite well, generating a corresponding high-resolution and high-quality texture which respects the character’s appearance description is not possible using current state of the art methods. Christian Murphy, Sudhir P. Mudur, Daniel Holden, Marc-André Carbonneau, Donya Ghafourzadeh, Andre Beauchamp |
MIG | 3 |
| 2020 | Learned motion matchingabstractIn this paper we present a learned alternative to the Motion Matching algorithm which retains the positive properties of Motion Matching but additionally achieves the scalability of neural-network-based generative models. Although neural-network-based generative models for character animation are capable of learning expressive, compact controllers from vast amounts of animation data, methods such as Motion Matching still remain a popular choice in the games industry due to their flexibility, predictability, low preprocessing time, and visual quality - all properties which can sometimes be difficult to achieve with neural-network-based methods. Yet, unlike neural networks, the memory usage of such methods generally scales linearly with the amount of data used, resulting in a constant trade-off between the diversity of animation which can be produced and real world production budgets. In this work we combine the benefits of both approaches and, by breaking down the Motion Matching algorithm into its individual steps, show how learned, scalable alternatives can be used to replace each operation in turn. Our final model has no need to store animation data or additional matching meta-data in memory, meaning it scales as well as existing generative models. At the same time, we preserve the behavior of Motion Matching, retaining the quality, control, and quick iteration time which are so important in the industry. Daniel Holden, Oussama Kanoun, Maksym Perepichka, Tiberiu Popa |
ACM Trans. Graph. | 1 |
| 2019 | Robust Marker Trajectory Repair for MOCAP using Kinematic ReferenceabstractProcessing motion capture data from optical markers for use in computer animations presents numerous technical challenges. Artifacts caused by noise, marker swaps, and marker occlusions often require manual intervention of a professionally trained marker tracking artist that spends large amounts of time and effort fixing these issues. Existing automatic solutions that attempt to fix marker data lack robustness due to either failing to properly detect and fix marker paths, or generating solutions that are challenging to integrate within current animation pipelines. In this paper, we present a method that robustly identifies invalid marker paths, removes the associated segments and generates new kinematically correct paths. We start by comparing the kinematic solutions generated by commercial software against the one generated by the state-of-the-art methods, using this information to determine which animation keyframes are invalid. Subsequently, we regenerate marker paths from the neural network based method [Holden 2018] and use a sophisticated marker filling algorithm to combine them with the original marker paths at sections where we detect the original data to be invalid. Our method outperforms alternatives by generating solutions that are both closer to the ground truth and more robust, allowing for manual intervention if required. Maksym Perepichka, Daniel Holden, Sudhir P. Mudur, Tiberiu Popa |
MIG | 2 |
| 2019 | DReCon: data-driven responsive control of physics-based charactersabstractInteractive control of self-balancing, physically simulated humanoids is a long standing problem in the field of real-time character animation. While physical simulation guarantees realistic interactions in the virtual world, simulated characters can appear unnatural if they perform unusual movements in order to maintain balance. Therefore, obtaining a high level of responsiveness to user control, runtime performance, and diversity has often been overlooked in exchange for motion quality. Recent work in the field of deep reinforcement learning has shown that training physically simulated characters to follow motion capture clips can yield high quality tracking results. We propose a two-step approach for building responsive simulated character controllers from unstructured motion capture data. First, meaningful features from the data such as movement direction, heading direction, speed, and locomotion style, are interactively specified and drive a kinematic character controller implemented using motion matching. Second, reinforcement learning is used to train a simulated character controller that is general enough to track the entire distribution of motion that can be generated by the kinematic controller. Our design emphasizes responsiveness to user input, visual quality, and low runtime cost for application in video-games. Kevin Bergamin, Simon Clavet, Daniel Holden, James Richard Forbes |
ACM Trans. Graph. | 3 |
| 2018 | Robust solving of optical motion capture data by denoisingabstractRaw optical motion capture data often includes errors such as occluded markers, mislabeled markers, and high frequency noise or jitter. Typically these errors must be fixed by hand - an extremely time-consuming and tedious task. Due to this, there is a large demand for tools or techniques which can alleviate this burden. In this research we present a tool that sidesteps this problem, and produces joint transforms directly from raw marker data (a task commonly called "solving") in a way that is extremely robust to errors in the input data using the machine learning technique of denoising. Starting with a set of marker configurations, and a large database of skeletal motion data such as the CMU motion capture database [CMU 2013b], we synthetically reconstruct marker locations using linear blend skinning and apply a unique noise function for corrupting this marker data - randomly removing and shifting markers to dynamically produce billions of examples of poses with errors similar to those found in real motion capture data. We then train a deep denoising feed-forward neural network to learn a mapping from this corrupted marker data to the corresponding transforms of the joints. Once trained, our neural network can be used as a replacement for the solving part of the motion capture pipeline, and, as it is very robust to errors, it completely removes the need for any manual clean-up of data. Our system is accurate enough to be used in production, generally achieving precision to within a few millimeters, while additionally being extremely fast to compute with low memory requirements. Daniel Holden |
ACM Trans. Graph. | 1 |
| 2017 | A Recurrent Variational Autoencoder for Human Motion Synthesis
Ikhsanul Habibie, Daniel Holden, Jonathan Schwarz, Joseph Yearsley, Taku Komura |
BMVC | 2 |
| 2017 | Phase-functioned neural networks for character controlabstractWe present a real-time character control mechanism using a novel neural network architecture called a Phase-Functioned Neural Network. In this network structure, the weights are computed via a cyclic function which uses the phase as an input. Along with the phase, our system takes as input user controls, the previous state of the character, the geometry of the scene, and automatically produces high quality motions that achieve the desired user control. The entire network is trained in an end-to-end fashion on a large dataset composed of locomotion such as walking, running, jumping, and climbing movements fitted into virtual environments. Our system can therefore automatically produce motions where the character adapts to different geometric environments such as walking and running over rough terrain, climbing over large rocks, jumping over obstacles, and crouching under low ceilings. Our network architecture produces higher quality results than time-series autoregressive models such as LSTMs as it deals explicitly with the latent variable of motion relating to the phase. Once trained, our system is also extremely fast and compact, requiring only milliseconds of execution time and a few megabytes of memory, even when trained on gigabytes of motion data. Our work is most appropriate for controlling characters in interactive scenes such as computer games and virtual reality systems. Daniel Holden, Taku Komura, Jun Saito |
ACM Trans. Graph. | 1 |
| 2017 | Learning Inverse Rig Mappings by Nonlinear RegressionabstractWe present a framework to design inverse rig-functions-functions that map low level representations of a character's pose such as joint positions or surface geometry to the representation used by animators called the animation rig. Animators design scenes using an animation rig, a framework widely adopted in animation production which allows animators to design character poses and geometry via intuitive parameters and interfaces. Yet most state-of-the-art computer animation techniques control characters through raw, low level representations such as joint angles, joint positions, or vertex coordinates. This difference often stops the adoption of state-of-the-art techniques in animation production. Our framework solves this issue by learning a mapping between the low level representations of the pose and the animation rig. We use nonlinear regression techniques, learning from example animation sequences designed by the animators. When new motions are provided in the skeleton space, the learned mapping is used to estimate the rig controls that reproduce such a motion. We introduce two nonlinear functions for producing such a mapping: Gaussian process regression and feedforward neural networks. The appropriate solution depends on the nature of the rig and the amount of data available for training. We show our framework applied to various examples including articulated biped characters, quadruped characters, facial animation rigs, and deformable characters. With our system, animators have the freedom to apply any motion synthesis algorithm to arbitrary rigging and animation pipelines for immediate editing. This greatly improves the productivity of 3D animation, while retaining the flexibility and creativity of artistic input. Daniel Holden, Jun Saito, Taku Komura |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2017 | Scanning and animating characters dressed in multiple-layer garments
Pengpeng Hu, Taku Komura, Daniel Holden, Yueqi Zhong |
Vis. Comput. | 3 |
| 2016 | A deep learning framework for character motion synthesis and editingabstractWe present a framework to synthesize character movements based on high level parameters, such that the produced movements respect the manifold of human motion, trained on a large motion capture dataset. The learned motion manifold, which is represented by the hidden units of a convolutional autoencoder, represents motion data in sparse components which can be combined to produce a wide range of complex movements. To map from high level parameters to the motion manifold, we stack a deep feedforward neural network on top of the trained autoencoder. This network is trained to produce realistic motion sequences from parameters such as a curve over the terrain that the character should follow, or a target location for punching and kicking. The feedforward control network and the motion manifold are trained independently, allowing the user to easily switch between feedforward networks according to the desired interface, without re-training the motion manifold. Once motion is generated it can be edited by performing optimization in the space of the motion manifold. This allows for imposing kinematic constraints, or transforming the style of the motion, while ensuring the edited motion remains natural. As a result, the system can produce smooth, high quality motion sequences without any manual pre-processing of the training data. Daniel Holden, Jun Saito, Taku Komura |
ACM Trans. Graph. | 1 |
| 2015 | Carpet unrolling for character control on uneven terrainabstractWe propose a type of relationship descriptor based on carpet unrolling that computes the joint positions of a character based on the sum of relative vectors originating from a local coordinate system embedded on the surface of a carpet. Given a terrain that a character is to walk over, the carpet is unrolled over the surface of the terrain. The carpet adapts to the geometry of the terrain and curves according to the trajectory of the character. Because trajectories of the body parts are computed as a weighted sum of the relative vectors, the character can smoothly adapt to the elevation of the terrain and the horizontal curves of the carpet. The carpet relationship descriptors are easy to parallelize and hundreds of characters can be animated in real-time by making use of the GPUs. This makes it applicable to real-time applications such as computer games. Mark Miller 0002, Daniel Holden, Rami Ali Al-Ashqar, Christophe Dubach, Kenny Mitchell, Taku Komura |
MIG | 2 |