VLDB 2026 Research / reviewers in the wild / expert
Darren Cosker
dblp:79/608 · also Darren P. Cosker
· DBLP profile ↗
51ranked-venue papers
2as first author
16since 2021 · last 2025
0000-0001-5177-4741ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 1 first-author · 14 since 2021Artificial intelligence and machine learning · 29 · 2 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | 3Diface: Synthesizing and Editing Holistic 3D Facial AnimationabstractCreating personalized 3D animations with precise control and realistic head motions remains challenging for current speech-driven 3D facial animation methods. Editing these animations is especially complex and time consuming, requires precise control and typically handled by highly skilled animators. Most existing works focus on controlling style or emotion of the synthesized animation and cannot edit/regenerate parts of an input animation. They also overlook the fact that multiple plausible lip and head movements can match the same audio input. To address these challenges, we present 3DiFACE, a novel method for holistic speech-driven 3D facial animation. Our approach produces diverse plausible lip and head motions for a single audio input and allows for editing via keyframing and interpolation. Specifically, we propose a fully-convolutional diffusion model that can leverage the viseme-level diversity in our training corpus. Additionally, we employ a speaking-style personalization and a novel sparsely-guided motion diffusion to enable precise control and editing. Through quantitative and qualitative evaluations, we demonstrate that our method is capable of generating and editing diverse holistic 3D facial animations given a single audio input, with control between high fidelity and diversity. Project page: https://balamuruganthambiraja.github.io/3DiFACE Balamurugan Thambiraja, Malte Prinzler, Mohammad Sadegh Ali Akbarian, Darren Cosker, Justus Thies |
3DV | 4 |
| 2025 | GASP: Gaussian Avatars with Synthetic PriorsabstractGaussian Splatting has changed the game for real-time photo-realistic rendering. One of the most popular applications of Gaussian Splatting is to create animatable avatars, known as Gaussian Avatars. Recent works have pushed the boundaries of quality and rendering efficiency but suffer from two main limitations. Either they require expensive multi-camera rigs to produce avatars with free-viewpoint rendering, or they can be trained with a single camera but only rendered at high quality from this fixed viewpoint. An ideal model would be trained using a short monocular video or image from available hardware, such as a webcam, and rendered from any view. To this end, we propose GASP: Gaussian Avatars with Synthetic Priors. To overcome the limitations of existing datasets, we exploit the pixel-perfect nature of synthetic data to train a Gaussian Avatar prior. By fitting this prior model to a single photo or video and fine-tuning it, we get a high-quality Gaussian Avatar, which supports 360° rendering. Our prior is only required for fitting, not inference, enabling real-time applications. Through our method, we obtain high-quality, animatable Avatars from limited data which can be animated and rendered at 70fps on commercial hardware. Jack R. Saunders, Charlie Hewitt, Yanan Jian, Marek Kowalski, Tadas Baltrusaitis, Yiye Chen, Darren Cosker, Virginia Estellers, Nicholas Gyde, Vinay P. Namboodiri, Ben Lundell |
CVPR | 7 |
| 2025 | The Honest Virtual Self? Effects of Avatar Personalization and Motor Control on Physiological Responses to Deceptive BehavioursabstractWith the advent of social virtual reality (VR), understanding how avatar personalization impacts users' social behaviour and physiological responses in VR is increasingly important. Deceptive behaviour is particularly relevant, as users can often hide their identity behind generic avatars, which can facilitate deception. We investigated whether avatar personalization and motor control impacted physiological reactions in a Detection-of-Deception task. Twenty participants performed the task with different levels of avatar personalization and motor control over the avatar. While personalization did not impact skin conductance responses, it led to a significant decrease in heart rate, which is an established physiological response associated with deceptive behaviour. This effect was exclusive to personalized avatars and not observed in generic avatars. Personalization and motor control led to increased embodiment, body ownership, agency and presence ratings. Overall, personalized avatars can preserve users’ physiological reactions to key social events, and thus enhance the realism of VR simulations. Anca Salagean, Darren Cosker, Danaë Emma Beckford Stanton Fraser |
ISMAR | 2 |
| 2024 | Dog Code: Human to Quadruped Embodiment using Shared CodebooksabstractMany VR animal embodiment sytsems suffer from poor animation fidelity, typically animating the animal avatars using inverse kinematics. We address this issue, presenting a novel deep-learning method, centred around a shared codebook, for mapping human motion to quadruped motion. Rather than trying to directly bridge the gap from human motion to quadruped motion, a task which has proven difficult, we first use a rule-based retargeter, relying on inverse and forward kinematics, to retarget human motions to an intermediate motion domain in which the motions share the same skeleton as the quadruped. We then use finite scalar quantization to construct a shared latent space, or codebook, between this intermediate domain and the quadruped motion domain. We do this by first pre-defining a finite number of discrete latent codes and then teaching these codes, using unsupervised deep-learning, to represent semantically similar motions in the two domains. We incorporate our real-time human-to-quadruped motion mapping into a VR quadruped embodiment system. The output quadruped animations are natural and realistic, while also preserving the semantics of users’ actions. Moreover, there is a strong synchrony between the input human motions and retargeted quadruped motions, an important factor for inducing a strong sense of VR embodiment. Dónal Egan, Alberto Jovane, Jan Szkaradek, George Fletcher 0002, Darren Cosker, Rachel McDonnell |
MIG | 5 |
| 2024 | RGBT-Dog: A Parametric Model and Pose Prior For Canine Body Analysis Data CreationabstractWhile there exists a great deal of labeled in-the-wild human data, the same is not true for animals. Manually creating new labels for the full range of animal species would take years of effort from the community. We are also now seeing the emerging potential for computer vision methods in areas like animal conservation, which is an additional motivation for this direction of research. Key to our approach is the ability to easily generate as many labeled training images as we desire across a range of different modalities. To achieve this, we present a new large scale canine motion capture dataset and parametric canine body and texture model. These are used to produce the first large scale, multi-domain, multi-task dataset for canine body analysis comprising of detailed synthetic labels on both real images and fully synthetic images in a range of realistic poses. We also introduce the first pose prior for animals in the form of a variational pose prior for canines which is used to fit the parametric model to images of canines. We demonstrate the effectiveness of our labels for training computer vision models on tasks such as parts-based segmentation and pose estimation and show such models can generalise to other animal species without additional training. Jake Deane, Sinead Kearney, Kwang In Kim, Darren Cosker |
WACV | 4 |
| 2024 | Look Ma, no markers: holistic performance capture without the hassleabstractWe tackle the problem of highly-accurate, holistic performance capture for the face, body and hands simultaneously. Motion-capture technologies used in film and game production typically focus only on face, body or hand capture independently, involve complex and expensive hardware and a high degree of manual intervention from skilled operators. While machine-learning-based approaches exist to overcome these problems, they usually only support a single camera, often operate on a single part of the body, do not produce precise world-space results, and rarely generalize outside specific contexts. In this work, we introduce the first technique for markerfree, high-quality reconstruction of the complete human body, including eyes and tongue, without requiring any calibration, manual intervention or custom hardware. Our approach produces stable world-space results from arbitrary camera rigs as well as supporting varied capture environments and clothing. We achieve this through a hybrid approach that leverages machine learning models trained exclusively on synthetic data and powerful parametric models of human shape and motion. We evaluate our method on a number of body, face and hand reconstruction benchmarks and demonstrate state-of-the-art results that generalize on diverse datasets. Charlie Hewitt, Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Lohit Petikam, Shideh Rezaeifar, Louis Florentin, Zafiirah Hosenie, Thomas J. Cashman 0001, Julien Valentin, Darren Cosker, Tadas Baltrusaitis |
ACM Trans. Graph. | 10 |
| 2024 | The Utilitarian Virtual Self - Using Embodied Personalized Avatars to Investigate Moral Decision-Making in Semi-Autonomous Vehicle DilemmasabstractEmbodied personalized avatars are a promising new tool to investigate moral decision-making by transposing the user into the "middle of the action" in moral dilemmas. Here, we tested whether avatar personalization and motor control could impact moral decision-making, physiological reactions and reaction times, as well as embodiment, presence and avatar perception. Seventeen participants, who had their personalized avatars created in a previous study, took part in a range of incongruent (i.e., harmful action led to better overall outcomes) and congruent (i.e., harmful action led to trivial outcomes) moral dilemmas as the drivers of a semi-autonomous car. They embodied four different avatars (counterbalanced - personalized motor control, personalized no motor control, generic motor control, generic no motor control). Overall, participants took a utilitarian approach by performing harmful actions only to maximize outcomes. We found increased physiological arousal (SCRs and heart rate) for personalized avatars compared to generic avatars, and increased SCRs in motor control conditions compared to no motor control. Participants had slower reaction times when they had motor control over their avatars, possibly hinting at more elaborate decision-making processes. Presence was also higher in motor control compared to no motor control conditions. Embodiment ratings were higher for personalized avatars, and generally, personalization and motor control were perceptually positive features. These findings highlight the utility of personalized avatars and open up a range of future research possibilities that could benefit from the affordances of this technology and simulate, more closely than ever, real-life action. Anca Salagean, Michelle Wu, George Fletcher 0002, Darren Cosker, Danaë Emma Beckford Stanton Fraser |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Meeting Your Virtual Twin: Effects of Photorealism and Personalization on Embodiment, Self-Identification and Perception of Self-Avatars in Virtual RealityabstractEmbodying virtual twins – photorealistic and personalized avatars – will soon be easily achievable in consumer-grade VR. For the first time, we explored how photorealism and personalization impact self-identification, as well as embodiment, avatar perception and presence. Twenty participants were individually scanned and, in a two-hour session, embodied four avatars (high photorealism personalized, low photorealism personalized, high photorealism generic, low photorealism generic). Questionnaire responses revealed stronger mid-immersion body ownership for the high photorealism personalized avatars compared to all other avatar types, and stronger embodiment for high photorealism compared to low photorealism avatars and for personalized compared to generic avatars. In a self-other face distinction task, participants took significantly longer to pause the face morphing videos of high photorealism personalized avatars, suggesting a stronger self-identification bias with these avatars. Photorealism and personalization were perceptually positive features; how employing these avatars in VR applications impacts users over time requires longitudinal investigation. Anca Salagean, Eleanor Crellin, Martin Parsons, Darren Cosker, Danaë Emma Beckford Stanton Fraser |
CHI | 4 |
| 2023 | HMD-NeMo: Online 3D Avatar Motion Generation From Sparse ObservationsabstractGenerating both plausible and accurate full body avatar motion is the key to the quality of immersive experiences in mixed reality scenarios. Head-Mounted Devices (HMDs) typically only provide a few input signals, such as head and hands 6-DoF. Recently, different approaches achieved impressive performance in generating full body motion given only head and hands signal. However, to the best of our knowledge, all existing approaches rely on full hand visibility. While this is the case when, e.g., using motion controllers, a considerable proportion of mixed reality experiences do not involve motion controllers and instead rely on egocentric hand tracking. This introduces the challenge of partial hand visibility owing to the restricted field of view of the HMD. In this paper, we propose the first unified approach, HMD-NeMo, that addresses plausible and accurate full body motion generation even when the hands may be only partially visible. HMD-NeMo is a lightweight neural network that predicts the full body motion in an online and real-time fashion. At the heart of HMD-NeMo is the spatiotemporal encoder with novel temporally adaptable mask tokens that encourage plausible motion in the absence of hand observations. We perform extensive analysis of the impact of different components in HMD-NeMo and introduce a new state-of-the-art on AMASS dataset through our evaluation. Mohammad Sadegh Ali Akbarian, Fatemehsadat Saleh, David Collier, Pashmina Cameron, Darren Cosker |
ICCV | 5 |
| 2023 | Imitator: Personalized Speech-driven 3D Facial AnimationabstractSpeech-driven 3D facial animation has been widely explored, with applications in gaming, character animation, virtual reality, and telepresence systems. State-of-the-art methods deform the face topology of the target actor to sync the input audio without considering the identity-specific speaking style and facial idiosyncrasies, thus, resulting in unrealistic and inaccurate lip movements. To address this, we present Imitator, a speech-driven facial expression synthesis method, which learns identity-specific details from a short input video and produces novel facial expressions matching the identity-specific speaking style and facial idiosyncrasies of the target actor. Specifically, we train a style-agnostic transformer on a large facial expression dataset which we use as a prior for audio-driven facial expressions. We utilize this prior to optimize for identity-specific speaking style based on a short reference video. To train the prior, we introduce a novel loss function based on detected bilabial consonants to ensure plausible lip closures and consequently improve the realism of the generated expressions. Through detailed experiments and user studies, we show that our approach improves Lip-Sync by 49% and produces expressive facial animations from input audio while preserving the actor’s speaking style. Project page: https://balamuruganthambiraja.github.io/Imitator Balamurugan Thambiraja, Ikhsanul Habibie, Mohammad Sadegh Ali Akbarian, Darren Cosker, Christian Theobalt, Justus Thies |
ICCV | 4 |
| 2023 | Probabilistic Human Mesh Recovery in 3D Scenes from Egocentric ViewsabstractAutomatic perception of human behaviors during social interactions is crucial for AR/VR applications, and an essential component is estimation of plausible 3D human pose and shape of our social partners from the egocentric view. One of the biggest challenges of this task is severe body truncation due to close social distances in egocentric scenarios, which brings large pose ambiguities for unseen body parts. To tackle this challenge, we propose a novel scene-conditioned diffusion method to model the body pose distribution. Conditioned on the 3D scene geometry, the diffusion model generates bodies in plausible humanscene interactions, with the sampling guided by a physics-based collision score to further resolve human-scene inter-penetrations. The classifier-free training enables flexible sampling with different conditions and enhanced diversity. A visibility-aware graph convolution model guided by per-joint visibility serves as the diffusion denoiser to incorporate inter-joint dependencies and per-body-part control. Extensive evaluations show that our method generates bodies in plausible interactions with 3D scenes, achieving both superior accuracy for visible joints and diversity for invisible body parts. The code is available at https://sanweiliti.github.io/egohmr/egohmr.html. Qianli Ma 0007, Yan Zhang 0054, Mohammad Sadegh Ali Akbarian, Darren Cosker, Siyu Tang 0001 |
ICCV | 5 |
| 2021 | Full-Body Motion from a Single Head-Mounted Device: Generating SMPL Poses from Partial ObservationsabstractThe increased availability and maturity of head-mounted and wearable devices opens up opportunities for remote communication and collaboration. However, the signal streams provided by these devices (e.g., head pose, hand pose, and gaze direction) do not represent a whole person. One of the main open problems is therefore how to leverage these signals to build faithful representations of the user. In this paper, we propose a method based on variational autoencoders to generate articulated poses of a human skeleton based on noisy streams of head and hand pose. Our approach relies on a model of pose likelihood that is novel and theoretically well-grounded. We demonstrate on publicly available datasets that our method is effective even from very impoverished signals and investigate how pose prediction can be made more accurate and realistic. Andrea Dittadi, Sebastian Dziadzio, Darren Cosker, Ben Lundell, Thomas J. Cashman 0001, Jamie Shotton |
ICCV | 3 |
| 2021 | How to train your dog: Neural enhancement of quadruped animationsabstractCreating realistic quadruped animations is challenging. Producing realistic animations using methods such as key-framing is time consuming and requires much artistic expertise. Alternatively, motion capture methods have their own challenges (getting the animal into a studio, attaching motion capture markers, and getting the animal to put on the desired performance) and the resulting animation will still most likely require cleaning up. It would be useful if an animator could provide an initial rough animation and in return be given a corresponding high quality realistic one. To this end, we present a deep-learning approach for the automatic enhancement of quadruped animations. Given an initial animation, possibly lacking the subtle details of true quadruped motion and/or containing small errors, our results show that it is possible for a neural network to learn how to add these subtleties and correct errors to produce an enhanced animation while preserving the semantics and context of the initial animation. Our work also has potential uses in other applications, for example, its ability to be used in real-time means it could form part of a quadruped embodiment system. Dónal Egan, George Fletcher 0002, Yiguo Qiao, Darren Cosker, Rachel McDonnell |
MIG | 4 |
| 2021 | Ego-Interaction: Visual Hand-Object Pose Correction for VR ExperiencesabstractImmersive virtual reality (VR) experiences may track both a user’s hands and a physical object at the same time and use the information to animate computer generated representations of the two interacting. However, to render visually without artefacts requires highly accurate tracking of the hands and the objects themselves as well as their relative locations – made even more difficult when the objects are articulated or deformable. If this tracking is incorrect, then the quality and immersion of the visual experience is reduced. In this paper we turn the problem around – instead of focusing on producing quality renders of hand-object interactions by improving tracking quality, we acknowledge there will be tracking errors and just focus on fixing the visualisations. We propose a Deep Neural Network (DNN) that modifies hand pose based on its relative position with the object. However, to train the network we require sufficient labelled data. We therefore also present a new dataset of hand-object interactions – Ego-Interaction. This is the first hand-object interaction dataset with egocentric RGBD videos and 3D ground truth data for both rigid and non-rigid objects. The Ego-Interaction dataset contains 92 sequences with 4 rigid, 1 articulated and 4 non-rigid objects and demonstrates hand-object interactions with 1 and 2 hands carefully captured, rigged and animated using motion capture. We provide our dataset as a general resource for researchers in the VR and AI community interested in other hand-object and egocentric tracking related problems. Catherine Taylor, Murray Evans, Eleanor Crellin, Martin Parsons, Darren Cosker |
MIG | 5 |
| 2021 | Fast, High-Quality Hierarchical Depth-Map Super-ResolutionabstractThe low spatial resolution of acquired depth maps is a major drawback of most RGBD sensors. However, there are many scenarios in which fast acquisition of high-resolution and high-quality depth maps would be desirable. One approach to achieve higher quality depth maps is through super-resolution. However, edge preservation is challenging, and artifacts such as depth confusion and blurring are easily introduced near boundaries. In view of this, we propose a method for fast, high-quality hierarchical depth-map super-resolution (HDS). In our method, a high-resolution RGB image is degraded layer by layer to guide the bilateral filtering of the depth map. To improve the upsampled depth map quality, we construct a feature-based bilateral filter (FBF) for the interpolation, by using the extracted RGB shallow and multi-layer features. To accelerate the process, we perform filtering only near depth boundaries and through matrix operations. We also propose an extension of our HDS model to a Classification-based Hierarchical Depth-map Super-resolution (C-HDS) model, where a context-aware trilateral filter reduces the contributions of unreliable neighbors to the current missing depth location. Experimental results show that the proposed method is significantly faster than existing methods for generating high-resolution depth maps, while also significantly improving depth quality compared to the current state-of-the-art approaches, especially for large-scale 16x super-resolution. Yiguo Qiao, Licheng Jiao, Wenbin Li 0002, Christian Richardt, Darren Cosker |
ACM Multimedia | 5 |
| 2021 | Automatic high fidelity foot contact location and timing for elite sprintingabstractAbstract Making accurate measurements of human body motions using only passive, non-interfering sensors such as video is a difficult task with a wide range of applications throughout biomechanics, health, sports and entertainment. The rise of machine learning-based human pose estimation has allowed for impressive performance gains, but machine learning-based systems require large datasets which might not be practical for niche applications. As such, it may be necessary to adapt systems trained for more general-purpose goals, but this might require a sacrifice in accuracy when compared with systems specifically developed for the application. This paper proposes two approaches to measuring a sprinter’s foot-ground contact locations and timing (step length and step frequency), a task which requires high accuracy. The first approach is a learning-free system based on occupancy maps. The second approach is a multi-camera 3D fusion of a state-of-the-art machine learning-based human pose estimation model. Both systems use the same underlying multi-camera system. The experiments show the learning-free computer vision algorithm to provide foot timing to better than 1 frame at 180 fps, and step length accurate to 7 mm, while the system based on pose estimation achieves timing better than 1.5 frames at 180 fps, and step length estimates accurate to 20 mm. Murray Evans, Steffi L. Colyer, Aki Salo, Darren Cosker |
Mach. Vis. Appl. | 4 |
| 2020 | RGBD-Dog: Predicting Canine Pose from RGBD SensorsabstractThe automatic extraction of animal 3D pose from images without markers is of interest in a range of scientific fields. Most work to date predicts animal pose from RGB images, based on 2D labelling of joint positions. However, due to the difficult nature of obtaining training data, no ground truth dataset of 3D animal motion is available to quantitatively evaluate these approaches. In addition, a lack of 3D animal pose data also makes it difficult to train 3D pose-prediction methods in a similar manner to the popular field of body-pose prediction. In our work, we focus on the problem of 3D canine pose estimation from RGBD images, recording a diverse range of dog breeds with several Microsoft Kinect v2s, simultaneously obtaining the 3D ground truth skeleton via a motion capture system. We generate a dataset of synthetic RGBD images from this data. A stacked hourglass network is trained to predict 3D joint locations, which is then constrained using prior models of shape and pose. We evaluate our model on both synthetic and real RGBD images and compare our results to previously published work fitting canine models to images. Finally, despite our training set consisting only of dog data, visual inspection implies that our network can produce good predictions for images of other quadrupeds - e.g. horses or cats - when their pose is similar to that contained in our training set. Sinead Kearney, Wenbin Li 0002, Martin Parsons, Kwang In Kim, Darren Cosker |
CVPR | 5 |
| 2020 | High-quality depth up-sampling via a supervised classification guided MRF model
Yiguo Qiao, Licheng Jiao, Xu Tang 0004, Wenbin Li 0002, Darren Cosker |
Pattern Recognit. Lett. | 5 |
| 2019 | Multi-character Motion Retargeting for Large-Scale Transformations
Maryam Naghizadeh, Darren Cosker |
CGI | 2 |
| 2019 | VR Props: An End-to-End Pipeline for Transporting Real Objects Into Virtual and Augmented EnvironmentsabstractImprovements in both software and hardware, as well as an increase in consumer suitable equipment, have resulted in great advances in the fields of virtual and augmented reality. Typically, systems use controllers or hand gestures to interact with virtual objects. However, these motions are often unnatural and diminish the immersion of the experience. Moreover, these approaches offer limited tactile feedback. There does not currently exist a platform to bring an arbitrary physical object into the virtual world without additional peripherals or the use of expensive motion capture systems. Such a system could be used for immersive experiences within the entertainment industry as well as being applied to VR or AR training experiences, in the fields of health and engineering. We propose an end-to-end pipeline for creating an interactive virtual prop from rigid and non-rigid physical objects. This includes a novel method for tracking the deformations of rigid and non-rigid objects at interactive rates using a single RGBD camera. We scan our physical object and process the point cloud to produce a triangular mesh. A range of possible deformations can be obtained by using a finite element method simulation and these are reduced to a low dimensional basis using principal component analysis. Machine learning approaches, in particular neural networks, have become key tools in computer vision and have been used on a range of tasks. Moreover, there has been an increased trend in training networks on synthetic data. To this end, we use a convolutional neural network, trained on synthetic data, to track the movement and potential deformations of an object in unlabelled RGB images from a single RGBD camera. We demonstrate our results for several objects with different sizes and appearances. Catherine Taylor, Chris Mullany, Robin McNicholas, Darren Cosker |
ISMAR | 4 |
| 2019 | User-Guided Facial Animation through an Evolutionary InterfaceabstractAbstract We propose a design framework to assist with user‐generated content in facial animation — without requiring any animation experience or ground truth reference. Where conventional prototyping methods rely on handcrafting by experienced animators, our approach looks to encode the role of the animator as an Evolutionary Algorithm acting on animation controls, driven by visual feedback from a user. Presented as a simple interface, users sample control combinations and select favourable results to influence later sampling. Over multiple iterations of disregarding unfavourable control values, parameters converge towards the user's ideal. We demonstrate our framework through two non‐trivial applications: creating highly nuanced expressions by evolving control values of a face rig and non‐linear motion through evolving control point positions of animation curves. Kyle Reed, Darren Cosker |
Comput. Graph. Forum | 2 |
| 2018 | Multi-Task Learning by Maximizing Statistical DependenceabstractWe present a new multi-task learning (MTL) approach that can be applied to multiple heterogeneous task estimators. Our motivation is that the best task estimator could change depending on the task itself. For example, we may have a deep neural network for the first task and a Gaussian process for the second task. Classical MTL approaches cannot handle this case, as they require the same model or even the same parameter types for all tasks. We tackle this by considering task-specific estimators as random variables. Then, the task relationships are discovered by measuring the statistical dependence between each pair of random variables. By doing so, our model is independent of the parametric nature of each task, and is even agnostic to the existence of such parametric formulation. We compare our algorithm with existing MTL approaches on challenging real world ranking and regression datasets, and show that our approach achieves comparable or better performance without knowing the parametric form. Youssef A. Mejjati, Darren Cosker, Kwang In Kim |
CVPR | 2 |
| 2018 | E-StopMotion: digitizing stop motion for enhanced animation and gamesabstractStop Motion Animation is the traditional craft of giving life to handmade models. The unique look and feel of this art form is hard to reproduce with 3D computer generated techniques. This is due to the unexpected details that appear from frame to frame and to the sometimes choppy appearance of the character movement. The artist's task can be overwhelming as he has to reshape a character into hundreds of poses to obtain just a few seconds of animation. The results of the animation are usually applied in 2D mediums like films or platform games. Character features that took a lot of effort to create thus remain unseen. We propose a novel system that allows the creation of 3D stop motion-like animations from 3D character shapes reconstructed from multi-view images. Given two or more reconstructed shapes from key frames, our method uses a combination of non-rigid registration and as-rigid-as-possible interpolation to generate plausible in-between shapes. This significantly reduces the artist's workload since much fewer poses are required. The reconstructed and interpolated shapes with complete 3D geometry can be manipulated even further through deformation techniques. The resulting shapes can then be used as animated characters in games or fused with 2D animation frames for enhanced stop motion films. Anamaria Ciucanu, Naval Bhandari, Xiaokun Wu 0001, Shridhar Ravikumar, Darren Cosker |
MIG | 6 |
| 2018 | Unsupervised Attention-guided Image-to-Image TranslationabstractCurrent unsupervised image-to-image translation techniques struggle to focus their attention on individual objects without altering the background or the way multiple objects interact within a scene. Motivated by the important role of attention in human perception, we tackle this limitation by introducing unsupervised attention mechanisms which are jointly adversarially trained with the generators and discriminators. We empirically demonstrate that our approach is able to attend to relevant regions in the image without requiring any additional supervision, and that by doing so it achieves more realistic mappings compared to recent approaches. Youssef A. Mejjati, Christian Richardt, James Tompkin 0001, Darren Cosker, Kwang In Kim |
NeurIPS | 4 |
| 2018 | Foot Contact Timings and Step Length for Sprint TrainingabstractThe frequency and length of a runner's steps are fundamental aspects of their performance. Accurate measurement of these parameters can provide valuable feedback to coaching staff, particularly if regular measurement can be made and monitored over the course of a season. This paper presents a computer vision based approach using high framerate cameras to measure the location and timing of foot contacts from which step length and frequency can be determined. The approach is evaluated against forceplates and optical motion capture for a mix of 18 trained and recreational runners. Force-plates and optical motion capture are considered to be the current "gold-standard" in biomechanics, and this is the first vision based paper to evaluate against these standards. Landing and take-off times were shown to be measurable to within 1.5 frames (at 180fps) and step length to within 1 cm. Murray Evans, Steffi L. Colyer, Darren Cosker, Aki Salo |
WACV | 3 |
| 2018 | Easy Generation of Facial Animation Using Motion GraphsabstractAbstract Facial animation is a time‐consuming and cumbersome task that requires years of experience and/or a complex and expensive set‐up. This becomes an issue, especially when animating the multitude of secondary characters required, e.g. in films or video‐games. We address this problem with a novel technique that relies on motion graphs to represent a landmarked database. Separate graphs are created for different facial regions, allowing a reduced memory footprint compared to the original data. The common poses are identified using a Euclidean‐based similarity metric and merged into the same node. This process traditionally requires a manually chosen threshold, however, we simplify it by optimizing for the desired graph compression. Motion synthesis occurs by traversing the graph using Dijkstra's algorithm, and coherent noise is introduced by swapping some path nodes with their neighbours. Expression labels, extracted from the database, provide the control mechanism for animation. We present a way of creating facial animation with reduced input that automatically controls timing and pose detail. Our technique easily fits within video‐game and crowd animation contexts, allowing the characters to be more expressive with less effort. Furthermore, it provides a starting point for content creators aiming to bring more life into their characters. José Serra, Ozan Cetinaslan, Shridhar Ravikumar, Verónica Orvalho, Darren Cosker |
Comput. Graph. Forum | 5 |
| 2018 | Learning system in real-time machine vision
Wenbin Li 0002, Zhihan Lyu, Darren Cosker |
Neurocomputing | 3 |
| 2018 | Learn to model blurry motion via directional similarity and filtering
Wenbin Li 0002, Da Chen 0003, Zhihan Lyu, Yan Yan 0002, Darren Cosker |
Pattern Recognit. | 5 |
| 2017 | Video interpolation using optical flow and Laplacian smoothness
Wenbin Li 0002, Darren Cosker |
Neurocomputing | 2 |
| 2017 | Blur robust optical flow using motion channel
Wenbin Li 0002, Jee Hang Lee, Gang Ren 0001, Darren Cosker |
Neurocomputing | 5 |
| 2017 | User-assisted image shadow removal
Han Gong, Darren Cosker |
Image Vis. Comput. | 2 |
| 2016 | User, metric, and computational evaluation of foveated rendering methodsabstractPerceptually lossless foveated rendering methods exploit human perception by selectively rendering at different quality levels based on eye gaze (at a lower computational cost) while still maintaining the user's perception of a full quality render. We consider three foveated rendering methods and propose practical rules of thumb for each method to achieve significant performance gains in real-time rendering frameworks. Additionally, we contribute a new metric for perceptual foveated rendering quality building on HDR-VDP2 that, unlike traditional metrics, considers the loss of fidelity in peripheral vision by lowering the contrast sensitivity of the model with visual eccentricity based on the Cortical Magnification Factor (CMF). The new metric is parameterized on user-test data generated in this study. Finally, we run our metric on a novel foveated rendering method for real-time immersive 360° content with motion parallax. Nicholas T. Swafford, José Antonio Iglesias Guitián, Charalampos Koniaris, Bochang Moon, Darren Cosker, Kenny Mitchell |
SAP | 5 |
| 2016 | Reading Between the Dots: Combining 3D Markers and FACS Classification for High-Quality Blendshape Facial Animation
Shridhar Ravikumar, Colin Davidson, Dmitry Kit, Neill D. F. Campbell, Luca Benedetti, Darren Cosker |
Graphics Interface | 6 |
| 2016 | Behavioural facial animation using motion graphs and mind mapsabstractWe present a new behavioural animation method that combines motion graphs for synthesis of animation and mind maps as behaviour controllers for the choice of motions, significantly reducing the cost of animating secondary characters. Motion graphs are created for each facial region from the analysis of a motion database, while synthesis occurs by minimizing the path distance that connects automatically chosen nodes. A Mind map is a hierarchical graph built on top of the motion graphs, where the user visually chooses how a stimulus affects the character's mood, which in turn will trigger motion synthesis. Different personality traits add more emotional complexity to the chosen reactions. Combining behaviour simulation and procedural animation leads to more emphatic and autonomous characters that react differently in each interaction, shifting the task of animating a character to one of defining its behaviour. José Serra, Verónica Orvalho, Darren Cosker |
MIG | 3 |
| 2016 | Fitting quadrics with a Bayesian priorabstractQuadrics are a compact mathematical formulation for a range of primitive surfaces. A problem arises when there are not enough data points to compute the model but knowledge of the shape is available. This paper presents a method for fitting a quadric with a Bayesian prior. We use a matrix normal prior in order to favour ellipsoids when fitting to ambiguous data. The results show the algorithm copes well when there are few points in the point cloud, competing with contemporary techniques in the area. Daniel Beale, Neill D. F. Campbell, Darren Cosker, Peter Hall 0001 |
Comput. Vis. Media | 4 |
| 2015 | Inferring changes in intrinsic parameters from motion blur
Alastair Barber, Matthew Brown 0001, Paul Hogbin, Darren Cosker |
Comput. Graph. | 4 |
| 2014 | Interactive Shadow Removal and Ground Truth for Variable Scene Categories
Han Gong, Darren Cosker |
BMVC | 2 |
| 2014 | Dual sensor filtering for robust tracking of head-mounted displaysabstractWe present a low-cost solution for yaw drift in head-mounted display systems that performs better than current commercial solutions and provides a wide capture area for pose tracking. Our method applies an extended Kalman filter to combine marker tracking data from an overhead camera with onboard head-mounted display accelerometer readings. To achieve low latency, we accelerate marker tracking with color blob localisation and perform this computation on the camera server, which only transmits essential pose data over WiFi for an unencumbered virtual reality system. Nicholas T. Swafford, Bas Boom, Kartic Subr, David Sinclair, Darren Cosker, Kenny Mitchell |
VRST | 5 |
| 2014 | Robust optical flow estimation for continuous blurred scenes using RGB-motion imaging and directional filteringabstractOptical flow estimation is a difficult task given real-world video footage with camera and object blur. In this paper, we combine a 3D pose&position tracker with an RGB sensor allowing us to capture video footage together with 3D camera motion. We show that the additional camera motion information can be embedded into a hybrid optical flow framework by interleaving an iterative blind deconvolution and warping based minimization scheme. Such a hybrid framework significantly improves the accuracy of optical flow estimation in scenes with strong blur. Our approach yields improved overall performance against three state-of-the-art baseline methods applied to our proposed ground truth sequences, as well as in several other real-world sequences captured by our novel imaging system. Wenbin Li 0002, Jee Hang Lee, Gang Ren 0001, Darren Cosker |
WACV | 5 |
| 2013 | Optical Flow Estimation Using Laplacian Mesh EnergyabstractIn this paper we present a novel non-rigid optical flow algorithm for dense image correspondence and non-rigid registration. The algorithm uses a unique Laplacian Mesh Energy term to encourage local smoothness whilst simul-taneously preserving non-rigid deformation. Laplacian de-formation approaches have become popular in graphics re-search as they enable mesh deformations to preserve local surface shape. In this work we propose a novel Laplacian Mesh Energy formula to ensure such sensible local defor-mations between image pairs. We express this wholly with-in the optical flow optimization, and show its application in a novel coarse-to-fine pyramidal approach. Our algorith-m achieves the state-of-the-art performance in all trials on the Garg et al. dataset, and top tier performance on the Middlebury evaluation. 1. Wenbin Li 0002, Darren Cosker, Matthew Brown 0001, Rui Tang 0015 |
CVPR | 2 |
| 2013 | User-aided single image shadow removalabstractThis paper presents a novel user-aided method for texture-preserving shadow removal from single images which only requires simple user input. Compared with the state-of-the-art, our algorithm addresses limitations in uneven shadow boundary processing and umbra recovery. We first detect an initial shadow boundary by growing a user specified shadow outline on an illumination-sensitive image. Interval-variable intensity sampling is introduced to avoid artefacts raised from uneven boundaries. We extract the initial scale field by applying local group intensity spline fittings around the shadow boundary. Bad intensity samples are replaced by their nearest alternatives based on a log-normal probability distribution of fitting errors. Finally, we use a gradual colour transfer to correct post-processing artefacts such as gamma correction and lossy compression. Compared with state-of-the-art methods, we offer highly user-friendly interaction, produce improved umbra recovery and improved processing given uneven shadow boundaries. Han Gong, Darren Cosker, Chuan Li 0001, Matthew Brown 0001 |
ICME | 2 |
| 2013 | Water Surface Modeling from a Single Viewpoint VideoabstractWe introduce a video-based approach for producing water surface models. Recent advances in this field output high-quality results but require dedicated capturing devices and only work in limited conditions. In contrast, our method achieves a good tradeoff between the visual quality and the production cost: It automatically produces a visually plausible animation using a single viewpoint video as the input. Our approach is based on two discoveries: first, shape from shading (SFS) is adequate to capture the appearance and dynamic behavior of the example water; second, shallow water model can be used to estimate a velocity field that produces complex surface dynamics. We will provide qualitative evaluation of our method and demonstrate its good performance across a wide range of scenes. Chuan Li 0001, David Pickup, Thomas Saunders, Darren Cosker, David Marshall 0001, Peter Hall 0001, Philip J. Willis |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2012 | An Anchor Patch Based Optimization Framework for Reducing Optical Flow Drift in Long Image Sequences
Wenbin Li 0002, Darren Cosker, Matthew Brown 0001 |
ACCV (3) | 2 |
| 2011 | A FACS valid 3D dynamic action unit database with applications to 3D dynamic morphable facial modelingabstractThis paper presents the first dynamic 3D FACS data set for facial expression research, containing 10 subjects performing between 19 and 97 different AUs both individually and in combination. In total the corpus contains 519 AU sequences. The peak expression frame of each sequence has been manually FACS coded by certified FACS experts. This provides a ground truth for 3D FACS based AU recognition systems. In order to use this data, we describe the first framework for building dynamic 3D morphable models. This includes a novel Active Appearance Model (AAM) based 3D facial registration and mesh correspondence scheme. The approach overcomes limitations in existing methods that require facial markers or are prone to optical flow drift. We provide the first quantitative assessment of such 3D facial mesh registration techniques and show how our proposed method provides more reliable correspondence. Darren Cosker, Eva Krumhuber, Adrian Hilton 0001 |
ICCV | 1 |
| 2010 | Reconstructing Mass-Conserved Water Surfaces Using Shape from Shading and Optical Flow
David Pickup, Chuan Li 0001, Darren Cosker, Peter Hall 0001, Philip J. Willis |
ACCV (4) | 3 |
| 2010 | Assessing the Uniqueness and Permanence of Facial Actions for Use in Biometric ApplicationsabstractAlthough the human face is commonly used as a physiological biometric, very little work has been done to exploit the idiosyncrasies of facial motions for person identification. In this paper, we investigate theuniquenessandpermanenceof facial actions to determine whether these can be used as a behavioral biometric. Experiments are carried out using 3-D video data of participants performing a set of very short verbal and nonverbal facial actions. The data have been collected over long time intervals to assess the variability of the subjects' emotional and physical conditions. Quantitative evaluations are performed for both the identification and the verification problems; the results indicate that emotional expressions (e.g., smile and disgust) are not sufficiently reliable for identity recognition in real-life situations, whereas speech-related facial movements show promising potential. Lanthao Benedikt, Darren Cosker, Paul L. Rosin, David Marshall 0001 |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2009 | Incremental Learning of Dynamical Models of FacesabstractActive Appearance Models (AAM) are a useful and popular tool for modelling facial variations. They have been used in face tracking, recognition and synthesis applications. For modelling facial dynamics of speech, they have been used in conjunction with Hid-den Markov Models (HMM). However, the high dimensionality of the training data and of the resulting AAMs leads to long learning time of HMMs and thus imposes serious limitations on their joint use. Here, we propose a new method for learning HMMs of facial dynamics incremen-tally. Our algorithm is fully unsupervised and can be used for on-line learning as new data becomes available. Another important feature of our algorithm is the automatic choice of the number of states in the model. We show in experiments an improvement in learning speed of three orders of magnitude. Finally, we demonstrate the quality of the learned HMMs by generating video footage of a talking face. 1 Introduction and Cyril Charron, Yulia Hicks, Peter Hall 0001, Darren Cosker |
BMVC | 4 |
| 2008 | Facial Dynamics in Biometric IdentificationabstractThis paper investigates the use of facial gestures for identity recognition. This is the first time that such a quantitative evaluation is conducted, comparing the analyses of 2D versus 3D dynamic data of verbal and nonverbal facial actions. Suitable data processing and feature extraction methods are examined, then a number of pattern matching techniques including the Fréchet distance, Correlation Coefficients, Hidden-Markov Models, Dynamic Time Warping and its derived forms are compared, in light of which an improved algorithm is proposed. Finally, a face recognition prototype using facial dynamics is built, achieving an Equal Error Rate EER=1.6%. 1 Lanthao Benedikt, Vedran Kajic, Darren Cosker, Paul L. Rosin, David Marshall 0001 |
BMVC | 3 |
| 2005 | Video assisted speech source separationabstractWe investigate the problem of integrating the complementary audio and visual modalities for speech separation. Rather than using independence criteria suggested in most blind source separation (BSS) systems, we use visual features from a video signal as additional information to optimize the unmixing matrix. We achieve this by using a statistical model characterizing the nonlinear coherence between audio and visual features as a separation criterion for both instantaneous and convolutive mixtures. We acquire the model by applying the Bayesian framework to the fused feature observations based on a training corpus. We point out several key existing challenges to the success of the system. Experimental results verify the proposed approach, which outperforms the audio only separation system in a noisy environment, and also provides a solution to the permutation problem. Wenwu Wang 0001, Darren Cosker, Yulia Hicks, Saeid Sanei, Jonathon A. Chambers |
ICASSP (5) | 2 |
| 2005 | Toward Perceptually Realistic Talking Heads: Models, Methods, and McGurkabstractMotivated by the need for an informative, unbiased, and quantitative perceptual method for the evaluation of a talking head we are developing, we propose a new test based on the “McGurk Effect.” Our approach helps to identify strengths and weaknesses in visual--speech synthesis algorithms for talking heads and facial animations, in general, and uses this insight to guide further development. We also evaluate the behavioral quality of our facial animations in comparison to real-speaker footage and demonstrate our tests by applying them to our current speech-driven facial animation system. Darren Cosker, David Marshall 0001, Paul L. Rosin, Susan Paddock, Simon K. Rushton |
ACM Trans. Appl. Percept. | 1 |
| 2001 | An expert system for multi-criteria decision making using Dempster Shafer theory
Malcolm J. Beynon, Darren Cosker, David Marshall 0001 |
Expert Syst. Appl. | 2 |