Kenny Mitchell

dblp:15/8317 · DBLP profile ↗
← Back
40ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0003-2420-7447ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 6 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 since 2021Artificial intelligence and machine learning · 6 · 2 since 2021
YearPublicationVenuePosition
2026 PAIR: A Pilot Dataset for Dual Perspective-based Video-Grounded Dialogue and Reconciliation
abstract
An ongoing challenge in task-based multi-agent systems is enabling collaborators to solve problems when each has only a partial view of the environment. Achieving a shared understanding requires more than just information exchange; it involves reconciling perspectives, negotiating, and engaging in joint problem-solving. This paper introduces PAIR, a pilot conversational corpus designed to capture how humans integrate complementary observations when interpreting dynamic video scenes. PAIR comprises 15 dialogues in which participants viewed the same event from egocentric and exocentric perspectives, then engaged in face-to-face discussions to construct a shared account. Transcripts were manually verified and annotated with 42 dialogue act categories, revealing interactional strategies such as questioning, clarification, and agreement. While lightweight, PAIR foregrounds collaborative sense-making in task-oriented dialogue, making it a controlled testbed for tasks such as dialogue act classification, video-grounded dialogue modelling, and multi-agent reasoning. Released openly, PAIR provides a foundation for developing and benchmarking systems that must reconcile complementary inputs to achieve common goals.
Lewis N. Watson, Carl Strathearn, Kenny Mitchell, Yanchao Yu
LREC3
2026 Frame2KG: A Benchmark and Evaluation Toolkit for Interpretable Frame-to-Graph Generation
abstract
This research focuses on interpretable frame-to-knowledge-graph (Frame2KG) generation for embodied robots, targeting on-device inference to enhance privacy, improve interpretability, and minimise compute costs. We introduce a novel corpus called Frame2KG-YC2, a synthetic, reproducible dataset, fine-tuned with Qwen2.5-VL models using LoRA on attention layers (QKVO), with and without GateProjUp/Down projections. For evaluation and benchmarking, we propose a deterministic toolkit featuring two-stage node matching (IoU gate + Hungarian assignment on text similarity) and comprehensive metrics. On a held-out test set, our best model achieves a Node F1μ=0.621, Edge F1μ=0.208, and a mean matched IoU of ≈0.61 with >98% schema conformity. Post-training quantisation maintains accuracy whilst promoting performance on edge hardware. We release the dataset, code, adapters, and evaluation toolkit to establish an interpretable baseline for future temporal and multi-view extensions.
Lewis N. Watson, Carl Strathearn, Kenny Mitchell, Yanchao Yu
LREC3
2025 NeFT-Net: N-window extended frequency transformer for rhythmic motion prediction
abstract
Advancements in prediction of human motion sequences are critical for enabling online virtual reality (VR) users to dance and move in ways that accurately mirror real-world actions, delivering a more immersive and connected experience. However, latency in networked motion tracking remains a significant challenge, disrupting engagement and necessitating predictive solutions to achieve real-time synchronization of remote motions. To address this issue, we propose a novel approach leveraging a synthetically generated dataset based on supervised foot anchor placement timings for rhythmic motions, ensuring periodicity and reducing prediction errors. Our model integrates a discrete cosine transform (DCT) to encode motion, refine high-frequency components, and smooth motion sequences, mitigating jittery artifacts. Additionally, we introduce a feed-forward attention mechanism designed to learn from N-window pairs of 3D key-point pose histories for precise future motion prediction. Quantitative and qualitative evaluations on the Human3.6M dataset highlight significant improvements in mean per joint position error (MPJPE) metrics, demonstrating the superiority of our technique over state-of-the-art approaches. We further introduce novel result pose visualizations through the use of generative AI methods.
Adeyemi Ademola, David Sinclair, Babis Koniaris, Samantha Hannah, Kenny Mitchell
Comput. Graph.5
2024 Expressive Talking Avatars
abstract
Stylized avatars are common virtual representations used in VR to support interaction and communication between remote collaborators. However, explicit expressions are notoriously difficult to create, mainly because most current methods rely on geometric markers and features modeled for human faces, not stylized avatar faces. To cope with the challenge of emotional and expressive generating talking avatars, we build the Emotional Talking Avatar Dataset which is a talking-face video corpus featuring 6 different stylized characters talking with 7 different emotions. Together with the dataset, we also release an emotional talking avatar generation method which enables the manipulation of emotion. We validated the effectiveness of our dataset and our method in generating audio based puppetry examples, including comparisons to state-of-the-art techniques and a user study. Finally, various applications of this method are discussed in the context of animating avatars in VR.
Shuai Tan 0002, Shengran Cheng, Qunfen Lin, Zijiao Zeng, Kenny Mitchell
IEEE Trans. Vis. Comput. Graph.6
2023 Real-time Facial Animation for 3D Stylized Character with Emotion Dynamics
abstract
Our aim is to improve animation production techniques' efficiency and effectiveness. We present two real-time solutions which drive character expressions in a geometrically consistent and perceptually valid way. Our first solution combines keyframe animation techniques with machine learning models. We propose a 3D emotion transfer network makes use of a 2D human image to generate a stylized 3D rig parameter. Our second solution combines blendshape-based motion capture animation techniques with machine learning models. We propose a blendshape adaption network which generates the character rig parameter motions with geometric consistency and temporally stability. We demonstrate the effectiveness of our system by comparing it to a commercial product Faceware. Results reveal that ratings of the recognition, intensity, and attractiveness of expressions depicted for animated characters via our systems are statistically higher than Faceware. Our results may be implemented into the animation pipeline, supporting animators to create expressions more rapidly and precisely.
Ruisi Zhang, Yu Ding 0001, Kenny Mitchell
ACM Multimedia5
2023 Emotional Voice Puppetry
abstract
The paper presents emotional voice puppetry, an audio-based facial animation approach to portray characters with vivid emotional changes. The lips motion and the surrounding facial areas are controlled by the contents of the audio, and the facial dynamics are established by category of the emotion and the intensity. Our approach is exclusive because it takes account of perceptual validity and geometry instead of pure geometric processes. Another highlight of our approach is the generalizability to multiple characters. The findings showed that training new secondary characters when the rig parameters are categorized as eye, eyebrows, nose, mouth, and signature wrinkles is significant in achieving better generalization results compared to joint training. User studies demonstrate the effectiveness of our approach both qualitatively and quantitatively. Our approach can be applicable in AR/VR and 3DUI, namely, virtual reality avatars/self-avatars, teleconferencing and in-game dialogue.
Ruisi Zhang, Shengran Cheng, Shuai Tan 0002, Yu Ding 0001, Kenny Mitchell, Xubo Yang
IEEE Trans. Vis. Comput. Graph.6
2023 Collimated Whole Volume Light Scattering in Homogeneous Finite Media
abstract
Crepuscular rays form when light encounters an optically thick or opaque medium which masks out portions of the visible scene. Real-time applications commonly estimate this phenomena by connecting paths between light sources and the camera after a single scattering event. We provide a set of algorithms for solving integration and sampling of single-scattered collimated light in a box-shaped medium and show how they extend to multiple scattering and convex media. First, a method for exactly integrating the unoccluded single scattering in rectilinear box-shaped medium is proposed and paired with a ratio estimator and moment-based approximation. Compared to previous methods, it requires only a single sample in unoccluded areas to compute the whole integral solution and provides greater convergence in the rest of the scene. Second, we derive an importance sampling scheme accounting for the entire geometry of the medium. This sampling strategy is then incorporated in an optimized Monte Carlo integration. The resulting integration scheme yields visible noise reduction and it is directly applicable to indoor scene rendering in room-scale interactive experiences. Furthermore, it extends to multiple light sources and achieves superior converge compared to independent sampling with existing algorithms. We validate our techniques against previous methods based on ray marching and distance sampling to prove their superior noise reduction capability.
Zdravko Velinov, Kenny Mitchell
IEEE Trans. Vis. Comput. Graph.2
2021 Foreword to the Special Section on the Reality-Virtuality Continuum and its Applications (RVCA)
Mashhuda Glencross, Kenny Mitchell, Mark Billinghurst
Comput. Graph.2
2021 Improving VIP viewer gaze estimation and engagement using adaptive dynamic anamorphosis
Kenny Mitchell
Int. J. Hum. Comput. Stud.2
2019 Photo-Realistic Facial Details Synthesis From Single Image
abstract
We present a single-image 3D face synthesis technique that can handle challenging facial expressions while recovering fine geometric details. Our technique employs expression analysis for proxy face geometry generation and combines supervised and unsupervised learning for facial detail synthesis. On proxy generation, we conduct emotion prediction to determine a new expression-informed proxy. On detail synthesis, we present a Deep Facial Detail Net (DFDN) based on Conditional Generative Adversarial Net (CGAN) that employs both geometry and appearance loss functions. For geometry, we capture 366 high-quality 3D scans from 122 different subjects under 3 facial expressions. For appearance, we use additional 163K in-the-wild face images and apply image-based rendering to accommodate lighting variations. Comprehensive experiments demonstrate that our framework can produce high-quality 3D faces with realistic details under challenging facial expressions.
Anpei Chen, Guli Zhang, Kenny Mitchell, Jingyi Yu 0001
ICCV4
2019 JUNGLE: An Interactive Visual Platform for Collaborative Creation and Consumption of Nonlinear Transmedia Stories
Mubbasir Kapadia, Carlos Muñiz 0001, Samuel S. Sohn, Sasha Schriber, Kenny Mitchell, Markus Gross 0001
ICIDS6
2019 Light Field Synthesis Using Inexpensive Surveillance Camera Systems
abstract
We present a light field synthesis technique that achieves accurate reconstruction given a low-cost, wide-baseline camera rig. Our system integrates optical flow with methods for rectification, disparity estimation, and feature extraction, which we then feed to a neural network view synthesis solver with wide-baseline capability. We propose two novel warping methods that improve the accuracy of disparity estimation and view synthesis. The methods enable the use of off-the-shelf surveillance camera hardware in a simplified and expedited capture workflow. A thorough analysis of the process and resulting view synthesis accuracy over state of the art is provided.
Frederike Dümbgen, Christopher Schroers, Kenny Mitchell
ICIP3
2019 Repurposing Labeled Photographs for Facial Tracking with Alternative Camera Intrinsics
abstract
Acquiring manually labeled training data for a specific application is expensive and while such data is often fully available for casual camera imagery, it is not a good fit for novel cameras. To overcome this, we present a repurposing approach that relies on spherical image warping to retarget an existing dataset of landmark labeled casual photography of people's faces with arbitrary poses from regular camera lenses to target cameras with significantly different intrinsics, such as those often attached to the head mounted displays (HMDs) with wide-angle lenses necessary to observe mouth and other features at close proximity and infrared only sensing for eye observations. Our method can predict landmarks of the HMD wearer in facial sub-regions in a divide-and-conquer fashion with particular focus on mouth and eyes. We demonstrate animated avatars in realtime using the face landmarks as input without user-specific nor application-specific dataset.
Caio José dos Santos Brito, Kenny Mitchell
VR2
2019 Compressed Animated Light Fields with Real-Time View-Dependent Reconstruction
abstract
We propose an end-to-end solution for presenting movie quality animated graphics to the user while still allowing the sense of presence afforded by free viewpoint head motion. By transforming offline rendered movie content into a novel immersive representation, we display the content in real-time according to the tracked head pose. For each frame, we generate a set of cubemap images per frame (colors and depths) using a sparse set of of cameras placed in the vicinity of the potential viewer locations. The cameras are placed with an optimization process so that the rendered data maximise coverage with minimum redundancy, depending on the lighting environment complexity. We compress the colors and depths separately, introducing an integrated spatial and temporal scheme tailored to high performance on GPUs for Virtual Reality applications. A view-dependent decompression algorithm decodes only the parts of the compressed video streams that are visible to users. We detail a real-time rendering algorithm using multi-view ray casting, with a variant that can handle strong view dependent effects such as mirror surfaces and glass. Compression rates of 150:1 and greater are demonstrated with quantitative analysis of image reconstruction quality and performance.
Charalampos Koniaris, Maggie Kosek, David Sinclair, Kenny Mitchell
IEEE Trans. Vis. Comput. Graph.4
2018 From Faces to Outdoor Light Probes
abstract
Abstract Image‐based lighting has allowed the creation of photo‐realistic computer‐generated content. However, it requires the accurate capture of the illumination conditions, a task neither easy nor intuitive, especially to the average digital photography enthusiast. This paper presents an approach to directly estimate an HDR light probe from a single LDR photograph, shot outdoors with a consumer camera, without specialized calibration targets or equipment. Our insight is to use a person's face as an outdoor light probe. To estimate HDR light probes from LDR faces we use an inverse rendering approach which employs data‐driven priors to guide the estimation of realistic, HDR lighting. We build compact, realistic representations of outdoor lighting both parametrically and in a data‐driven way, by training a deep convolutional autoencoder on a large dataset of HDR sky environment maps. Our approach can recover high‐frequency, extremely high dynamic range lighting environments. For quantitative evaluation of lighting estimation accuracy and relighting accuracy, we also contribute a new database of face photographs with corresponding HDR light probes. We show that relighting objects with HDR light probes estimated by our method yields realistic results in a wide variety of settings.
Dan Andrei Calian, Jean-François Lalonde, Paulo F. U. Gotardo, Tomas Simon, Iain A. Matthews, Kenny Mitchell
Comput. Graph. Forum6
2018 Empowerment and embodiment for collaborative mixed reality systems
abstract
Abstract We present several mixed‐reality‐based remote collaboration settings by using consumer head‐mounted displays. We investigated how two people are able to work together in these settings. We found that the person in the AR system will be regarded as the “leader” (i.e., they provide a greater contribution to the collaboration), whereas no similar “leader” emerges in augmented reality (AR)‐to‐AR and AR‐to‐VRBody settings. We also found that these special patterns of leadership only emerged for 3D interactions and not for 2D interactions. Results about the participants' experience of leadership, collaboration, embodiment, presence, and copresence shed further light on these findings.
David Sinclair, Kenny Mitchell
Comput. Animat. Virtual Worlds3
2017 Real-time Rendering with Compressed Animated Light Fields
Babis Koniaris, Maggie Kosek, David Sinclair, Kenny Mitchell
Graphics Interface4
2017 Rapid one-shot acquisition of dynamic VR avatars
abstract
We present a system for rapid acquisition of bespoke, animatable, full-body avatars including face texture and shape. A blendshape rig with a skeleton is used as a template for customization. Identity blendshapes are used to customize the body and face shape at the fitting stage, while animation blendshapes allow the face to be animated. The subject assumes a T-pose and a single snapshot is captured using a stereo RGB plus depth sensor rig. Our system automatically aligns a photo texture and fits the 3D shape of the face. The body shape is stylized according to body dimensions estimated from segmented depth. The face identity blendweights are optimised according to image-based facial landmarks, while a custom texture map for the face is generated by warping the input images to a reference texture according to the facial landmarks. The total capture and processing time is under 10 seconds and the output is a light-weight, game-engine-ready avatar which is recognizable as the subject. We demonstrate our system in a VR environment in which each user sees the other users' animated avatars through a VR headset with real-time audio-based facial animation and live body motion tracking, affording an enhanced level of presence and social engagement compared to generic avatars.
Charles Malleson, Maggie Kosek, Martin Klaudiny, Ivan Huerta Casado, Jean-Charles Bazin, Alexander Sorkine-Hornung, Mark Mine, Kenny Mitchell
VR8
2017 Demonstration: Rapid one-shot acquisition of dynamic VR avatars
abstract
In this demonstration, we showcase a system for rapid acquisition of bespoke avatars for each participant (subject) in a social VR environment is presented. For each subject, the system automatically customizes a parametric avatar model to match the captured subject by adjusting its overall height, body and face shape parameters and generating a custom face texture.
Charles Malleson, Maggie Kosek, Martin Klaudiny, Ivan Huerta Casado, Jean-Charles Bazin, Alexander Sorkine-Hornung, Mark Mine, Kenny Mitchell
VR8
2017 Real-Time Multi-View Facial Capture with Synthetic Training
abstract
We present a real-time multi-view facial capture system facilitated by synthetic training imagery. Our method is able to achieve high-quality markerless facial performance capture in real-time from multi-view helmet camera data, employing an actor specific regressor. The regressor training is tailored to specified actor appearance and we further condition it for the expected illumination conditions and the physical capture rig by generating the training data synthetically. In order to leverage the information present in live imagery, which is typically provided by multiple cameras, we propose a novel multi-view regression algorithm that uses multi-dimensional random ferns. We show that higher quality can be achieved by regressing on multiple video streams than previous approaches that were designed to operate on only a single view. Furthermore, we evaluate possible camera placements and propose a novel camera configuration that allows to mount cameras outside the field of view of the actor, which is very beneficial as the cameras are then less of a distraction for the actor and allow for an unobstructed line of sight to the director and other actors. Our new real-time facial capture approach has immediate application in on-set virtual production, in particular with the ever-growing demand for motion-captured facial animation in visual effects and video games.
Martin Klaudiny, Steven McDonagh 0001, Derek Bradley, Thabo Beeler, Kenny Mitchell
Comput. Graph. Forum5
2017 Noise Reduction on G-Buffers for Monte Carlo Filtering
abstract
Abstract We propose a novel pre‐filtering method that reduces the noise introduced by depth‐of‐field and motion blur effects in geometric buffers (G‐buffers) such as texture, normal and depth images. Our pre‐filtering uses world positions and their variances to effectively remove high‐frequency noise while carefully preserving high‐frequency edges in the G‐buffers. We design a new anisotropic filter based on a per‐pixel covariance matrix of world position samples. A general error estimator, Stein's unbiased risk estimator, is then applied to estimate the optimal trade‐off between the bias and variance of pre‐filtered results. We have demonstrated that our pre‐filtering improves the results of existing filtering methods numerically and visually for challenging scenes where depth‐of‐field and motion blurring introduce a significant amount of noise in the G‐buffers.
Bochang Moon, José Antonio Iglesias Guitián, Steven McDonagh 0001, Kenny Mitchell
Comput. Graph. Forum4
2016 Synthetic Prior Design for Real-Time Face Tracking
abstract
Real-time facial performance capture has recently been gaining popularity in virtual film production, driven by advances in machine learning, which allows for fast inference of facial geometry from video streams. These learning-based approaches are significantly influenced by the quality and amount of labelled training data. Tedious construction of training sets from real imagery can be replaced by rendering a facial animation rig under on-set conditions expected at runtime. We learn a synthetic actor-specific prior by adapting a state-of-the-art facial tracking method. Synthetic training significantly reduces the capture and annotation burden and in theory allows generation of an arbitrary amount of data. But practical realities such as training time and compute resources still limit the size of any training set. We construct better and smaller training sets by investigating which facial image appearances are crucial for tracking accuracy, covering the dimensions of expression, viewpoint and illumination. A reduction of training data in 1-2 orders of magnitude is demonstrated whilst tracking accuracy is retained for challenging on-set footage.
Steven McDonagh 0001, Martin Klaudiny, Derek Bradley, Thabo Beeler, Iain A. Matthews, Kenny Mitchell
3DV6
2016 User, metric, and computational evaluation of foveated rendering methods
abstract
Perceptually lossless foveated rendering methods exploit human perception by selectively rendering at different quality levels based on eye gaze (at a lower computational cost) while still maintaining the user's perception of a full quality render. We consider three foveated rendering methods and propose practical rules of thumb for each method to achieve significant performance gains in real-time rendering frameworks. Additionally, we contribute a new metric for perceptual foveated rendering quality building on HDR-VDP2 that, unlike traditional metrics, considers the loss of fidelity in peripheral vision by lowering the contrast sensitivity of the model with visual eccentricity based on the Cortical Magnification Factor (CMF). The new metric is parameterized on user-test data generated in this study. Finally, we run our metric on a novel foveated rendering method for real-time immersive 360° content with motion parallax.
Nicholas T. Swafford, José Antonio Iglesias Guitián, Charalampos Koniaris, Bochang Moon, Darren Cosker, Kenny Mitchell
SAP6
2016 Nonlinearly Weighted First-order Regression for Denoising Monte Carlo Renderings
abstract
We address the problem of denoising Monte Carlo renderings by studying existing approaches and proposing a new algorithm that yields state-of-the-art performance on a wide range of scenes. We analyze existing approaches from a theoretical and empirical point of view, relating the strengths and limitations of their corresponding components with an emphasis on production requirements. The observations of our analysis instruct the design of our new filter that offers high-quality results and stable performance. A key observation of our analysis is that using auxiliary buffers (normal, albedo, etc.) to compute the regression weights greatly improves the robustness of zero-order models, but can be detrimental to first-order models. Consequently, our filter performs a first-order regression leveraging a rich set of auxiliary buffers only when fitting the data, and, unlike recent works, considers the pixel color alone when computing the regression weights. We further improve the quality of our output by using a collaborative denoising scheme. Lastly, we introduce a general mean squared error estimator, which can handle the collaborative nature of our filter and its nonlinear weights, to automatically set the bandwidth of our regression kernel.
Benedikt Bitterli, Fabrice Rousselle, Bochang Moon, José Antonio Iglesias Guitián, David Adler, Kenny Mitchell, Wojciech Jarosz, Jan Novák
Comput. Graph. Forum6
2016 Pixel History Linear Models for Real-Time Temporal Filtering
abstract
Abstract We propose a new real‐time temporal filtering and antialiasing (AA) method for rasterization graphics pipelines. Our method is based on Pixel History Linear Models (PHLM), a new concept for modeling the history of pixel shading values over time using linear models. Based on PHLM, our method can predict per‐pixel variations of the shading function between consecutive frames. This combines temporal reprojection with per‐pixel shading predictions in order to provide temporally coherent shading, even in the presence of very noisy input images. Our method can address both spatial and temporal aliasing problems under a unique filtering framework that minimizes filtering error through a recursive least squares algorithm. We demonstrate our method working with a commercial deferred shading engine for rasterization and with our own OpenGL deferred shading renderer. We have implemented our method in GPU and it has shown significant reduction of temporal flicker in very challenging scenarios including foliage rendering, complex non‐linear camera motions, dynamic lighting, reflections, shadows and fine geometric details. Our approach, based on PHLM, avoids the creation of visible ghosting artifacts and it reduces the filtering overblur characteristic of temporal deflickering methods. At the same time, the results are comparable to state‐of‐the‐art real‐time filters in terms of temporal coherence.
José Antonio Iglesias Guitián, Bochang Moon, Charalampos Koniaris, Eric Smolikowski, Kenny Mitchell
Comput. Graph. Forum5
2016 Adaptive polynomial rendering
abstract
In this paper, we propose a new adaptive rendering method to improve the performance of Monte Carlo ray tracing, by reducing noise contained in rendered images while preserving high-frequency edges. Our method locally approximates an image with polynomial functions and the optimal order of each polynomial function is estimated so that our reconstruction error can be minimized. To robustly estimate the optimal order, we propose a multi-stage error estimation process that iteratively estimates our reconstruction error. In addition, we present an energy-preserving outlier removal technique to remove spike noise without causing noticeable energy loss in our reconstruction result. Also, we adaptively allocate additional ray samples to high error regions guided by our error estimation. We demonstrate that our approach outperforms state-of-the-art methods by controlling the tradeoff between reconstruction bias and variance through locally defining our polynomial order, even without need for filtering bandwidth optimization, the common approach of other recent methods.
Bochang Moon, Steven McDonagh 0001, Kenny Mitchell, Markus Gross 0001
ACM Trans. Graph.3
2015 Online view sampling for estimating depth from light fields
abstract
Geometric information such as depth obtained from light fields finds more applications recently. Where and how to sample images to populate a light field is an important problem to maximize the usability of information gathered for depth reconstruction. We propose a simple analysis model for view sampling and an adaptive, online sampling algorithm tailored to light field depth reconstruction. Our model is based on the trade-off between visibility and depth resolvability for varying sampling locations, and seeks the optimal locations that best balance the two conflicting criteria.
Changil Kim 0001, Kartic Subr, Kenny Mitchell, Alexander Sorkine-Hornung, Markus Gross 0001
ICIP3
2015 Carpet unrolling for character control on uneven terrain
abstract
We propose a type of relationship descriptor based on carpet unrolling that computes the joint positions of a character based on the sum of relative vectors originating from a local coordinate system embedded on the surface of a carpet. Given a terrain that a character is to walk over, the carpet is unrolled over the surface of the terrain. The carpet adapts to the geometry of the terrain and curves according to the trajectory of the character. Because trajectories of the body parts are computed as a weighted sum of the relative vectors, the character can smoothly adapt to the elevation of the terrain and the horizontal curves of the carpet. The carpet relationship descriptors are easy to parallelize and hundreds of characters can be animated in real-time by making use of the GPUs. This makes it applicable to real-time applications such as computer games.
Mark Miller 0002, Daniel Holden, Rami Ali Al-Ashqar, Christophe Dubach, Kenny Mitchell, Taku Komura
MIG5
2015 Adaptive rendering with linear predictions
abstract
We propose a new adaptive rendering algorithm that enhances the performance of Monte Carlo ray tracing by reducing the noise, i.e., variance, while preserving a variety of high-frequency edges in rendered images through a novel prediction based reconstruction. To achieve our goal, we iteratively build multiple, but sparse linear models. Each linear model has its prediction window, where the linear model predicts the unknown ground truth image that can be generated with an infinite number of samples. Our method recursively estimates prediction errors introduced by linear predictions performed with different prediction windows, and selects an optimal prediction window minimizing the error for each linear model. Since each linear model predicts multiple pixels within its optimal prediction interval, we can construct our linear models only at a sparse set of pixels in the image screen. Predicting multiple pixels with a single linear model poses technical challenges, related to deriving error analysis for regions rather than pixels, and has not been addressed in the field. We address these technical challenges, and our method with robust error analysis leads to a drastically reduced reconstruction time even with higher rendering quality, compared to state-of-the-art adaptive methods. We have demonstrated that our method outperforms previous methods numerically and visually with high performance ray tracing kernels such as OptiX and Embree.
Bochang Moon, José Antonio Iglesias Guitián, Sung-Eui Yoon, Kenny Mitchell
ACM Trans. Graph.4
2014 Influence of animated reality mixing techniques on user experience
abstract
We investigate the influence of motion effects in the domain of mobile Augmented Reality (AR) games on user experience and task performance. The work focuses on evaluating responses to a selection of synthesized camera oriented reality mixing techniques for AR, such as motion blur, defocus blur, latency and lighting responsiveness. In our cross section of experiments, we observe that these measures have a significant impact on perceived realism, where aesthetic quality is valued. However, lower latency records the strongest correlation with improved subjective enjoyment, satisfaction, and realism, and objective scoring performance. We conclude that the reality mixing techniques employed are not significant in the overall user experience of a mobile AR game, except where harmonious or convincing blended AR image quality is consciously desired by the participants.
Fabio Zünd, Marcel Lancelle, Mattia Ryffel, Robert W. Sumner, Kenny Mitchell, Markus Gross 0001
MIG5
2014 Poxels: polygonal voxel environment rendering
abstract
We present efficient rendering of opaque, sparse, voxel environments with data amplified in local graphics memory with stream-out from a geomery shader to a cached vertex buffer pool. We show that our Poxel rendering primitive aligns with optimized rasterization hardware and so results in high visual quality over ray casting methods. Lossless run length encoding of occlusion culled voxels and coordinate quantization further reduces host data transfers.
Mark Miller 0002, Andrew Cumming, Kevin Chalmers, Benjamin Kenwright, Kenny Mitchell
VRST5
2014 Dual sensor filtering for robust tracking of head-mounted displays
abstract
We present a low-cost solution for yaw drift in head-mounted display systems that performs better than current commercial solutions and provides a wide capture area for pose tracking. Our method applies an extended Kalman filter to combine marker tracking data from an overhead camera with onboard head-mounted display accelerometer readings. To achieve low latency, we accelerate marker tracking with color blob localisation and perform this computation on the camera server, which only transmits essential pose data over WiFi for an unencumbered virtual reality system.
Nicholas T. Swafford, Bas Boom, Kartic Subr, David Sinclair, Darren Cosker, Kenny Mitchell
VRST6
2014 Visibility Silhouettes for Semi-Analytic Spherical Integration
abstract
Abstract At each shade point, the spherical visibility function encodes occlusion from surrounding geometry, in all directions. Computing this function is difficult and point‐sampling approaches, such as ray‐tracing or hardware shadow mapping, are traditionally used to efficiently approximate it. We propose a semi‐analytic solution to the problem where the spherical silhouette of the visibility is computed using a search over a 4D dual mesh of the scene. Once computed, we are able to semi‐analytically integrate visibility‐masked spherical functions along the visibility silhouette, instead of over the entire hemisphere. In this way, we avoid the artefacts that arise from using point‐sampling strategies to integrate visibility, a function with unbounded frequency content. We demonstrate our approach on several applications, including direct illumination from realistic lighting and computation of pre‐computed radiance transfer data. Additionally, we present a new frequency‐space method for exactly computing all‐frequency shadows on diffuse surfaces. Our results match ground truth computed using importance‐sampled stratified Monte Carlo ray‐tracing, with comparable performance on scenes with low‐to‐moderate geometric complexity.
Derek Nowrouzezahrai, Ilya Baran, Kenny Mitchell, Wojciech Jarosz
Comput. Graph. Forum3
2014 Error analysis of estimators that use combinations of stochastic sampling strategies for direct illumination
abstract
Abstract We present a theoretical analysis of error of combinations of Monte Carlo estimators used in image synthesis. Importance sampling and multiple importance sampling are popular variance‐reduction strategies. Unfortunately, neither strategy improves the rate of convergence of Monte Carlo integration. Jittered sampling (a type of stratified sampling), on the other hand is known to improve the convergence rate. Most rendering software optimistically combine importance sampling with jittered sampling, hoping to achieve both. We derive the exact error of the combination of multiple importance sampling with jittered sampling. In addition, we demonstrate a further benefit of introducing negative correlations (antithetic sampling) between estimates to the convergence rate. As with importance sampling, antithetic sampling is known to reduce error for certain classes of integrands without affecting the convergence rate. In this paper, our analysis and experiments reveal that importance and antithetic sampling, if used judiciously and in conjunction with jittered sampling, may improve convergence rates. We show the impact of such combinations of strategies on the convergence rate of estimators for direct illumination.
Kartic Subr, Derek Nowrouzezahrai, Wojciech Jarosz, Jan Kautz, Kenny Mitchell
Comput. Graph. Forum5
2013 Multi-spectral Material Classification in Landscape Scenes Using Commodity Hardware
Gwyneth Bradbury, Kenny Mitchell, Tim Weyrich
CAIP (2)2
2013 Simulated motion blur does not improve player experience in racing game
abstract
Motion blur effects are commonly used in racing games [Sousa 2008; Vlachos 2008; Ritchie et al. 2010] to add a sense of realism as well as to minimize artifacts due to strobing and temporal aliasing [Glassner 1999]. Typically, motion blur computations are expensive, and for real-time applications, trade-offs are made between the quality of the effects and the computational cost. In this work, we wanted to understand: (i) the practical impact of the motion blur effect on the player experience; and (ii) whether the value gained by including the effect is worth the extra cost in computation, real-time performance, development time, etc. We studied the objective and subjective aspects of the player experience for Split Second: Velocity (Black Rock Studios, Disney), a high-speed racing game, in the presence and absence of the motion blur effect. We found that neither objective measures of participants' performance (e.g., time to complete a race) nor subjective measures of the player experience (e.g, enjoyment of a race, perceived speed) were affected, even though participants could reliably detect the presence of the motion blur effect. We conclude that motion blur effects, while useful for reducing artifacts and achieving a realistic 'look', do not significantly enhance the player experience.
Lavanya Sharan, Zhe Han Neo, Kenny Mitchell, Jessica K. Hodgins
MIG3
2012 Iterative Image Warping
abstract
Abstract Animated image sequences often exhibit a large amount of inter‐frame coherence which standard rendering algorithms and pipelines are ill‐equipped to exploit, limiting their efficiency. To address this inefficiency we transfer rendering results across frames using a novel image warping algorithm based on fixed point iteration. We analyze the behavior of the iteration and describe two alternative algorithms designed to suit different performance requirements. Further, to demonstrate the versatility of our approach we apply it to a number of spatio‐temporal rendering problems including 30‐to‐60Hz frame upsampling, stereoscopic 3D conversion, defocus and motion blur. Finally we compare our approach against existing image warping methods and demonstrate a significant performance improvement.
Huw Bowles, Kenny Mitchell, Robert W. Sumner, Jeremy Moore, Markus Gross 0001
Comput. Graph. Forum2
2011 Light factorization for mixed-frequency shadows in augmented reality
abstract
Integrating animated virtual objects with their surroundings for high-quality augmented reality requires both geometric and radio-metric consistency. We focus on the latter of these problems and present an approach that captures and factorizes external lighting in a manner that allows for realistic relighting of both animated and static virtual objects. Our factorization facilitates a combination of hard and soft shadows, with high-performance, in a manner that is consistent with the surrounding scene lighting.
Derek Nowrouzezahrai, Stefan Geiger, Kenny Mitchell, Robert W. Sumner, Wojciech Jarosz, Markus Gross 0001
ISMAR3
2011 Modular Radiance Transfer
abstract
Many rendering algorithms willingly sacrifice accuracy, favoring plausible shading with high-performance. Modular Radiance Transfer (MRT) models coarse-scale, distant indirect lighting effects in scene geometry that scales from high-end GPUs to low-end mobile platforms. MRT eliminates scene-dependent precomputation by storing compact transport on simple shapes, akin to bounce cards used in film production. These shapes' modular transport can be instanced, warped and connected on-the-fly to yield approximate light transport in large scenes. We introduce a prior on incident lighting distributions and perform all computations in low-dimensional subspaces. An implicit lighting environment induced from the low-rank approximations is in turn used to model secondary effects, such as volumetric transport variation, higher-order irradiance, and transport through lightfields. MRT is a new approach to precomputed lighting that uses a novel low-dimensional subspace simulation of light transport to uniquely balance the need for high-performance and portable solutions, low memory usage, and fast authoring iteration.
Brad Loos, Lakulish Antani, Kenny Mitchell, Derek Nowrouzezahrai, Wojciech Jarosz, Peter-Pike J. Sloan
ACM Trans. Graph.3
2011 OSCAM - optimized stereoscopic camera control for interactive 3D
abstract
This paper presents a controller for camera convergence and interaxial separation that specifically addresses challenges ininteractivestereoscopic applications like games. In such applications, unpredictable viewer- or object-motion often compromises stereopsis due to excessive binocular disparities. We derive constraints on the camera separation and convergence that enable our controller to automatically adapt to any given viewing situation and 3D scene, providing an exact mapping of the virtual content into a comfortable depth range around the display. Moreover, we introduce an interpolation function that linearizes the transformation of stereoscopic depth over time, minimizing nonlinear visual distortions. We describe how to implement the complete control mechanism on the GPU to achieve running times below 0.2ms for full HD. This provides a practical solution even for demanding real-time applications. Results of a user study show a significant increase of stereoscopic comfort, without compromising perceived realism. Our controller enables 'fail-safe' stereopsis, provides intuitive control to accommodate to personal preferences, and allows to properly display stereoscopic content on differently sized output devices.
Thomas Oskam, Alexander Sorkine-Hornung, Huw Bowles, Kenny Mitchell, Markus Gross 0001
ACM Trans. Graph.4