VLDB 2026 Research / reviewers in the wild / expert
John P. Lewis
dblp:83/6305 · also J. P. Lewis 0001, John Lewis 0001, John Peter Lewis
· DBLP profile ↗
51ranked-venue papers
12as first author
6since 2021 · last 2025
0000-0002-6835-7263ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 45 · 11 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 16 · 8 first-author · 1 since 2021Artificial intelligence and machine learning · 14 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On Exact Bit-level Reversible Transformers Without Changing ArchitectureabstractIn this work we present the BDIA-transformer, which is an exact bit-level reversible transformer that uses an unchanged standard architecture for inference. The basic idea is to first treat each transformer block as the Euler integration approximation for solving an ordinary differential equation (ODE) and then incorporate the technique of bidirectional integration approximation (BDIA) (originally designed for diffusion inversion) into the neural architecture, together with activation quantization to make it exactly bit-level reversible. In the training process, we let a hyper-parameter $\gamma$ in BDIA-transformer randomly take one of the two values $\{0.5, -0.5\}$ per training sample per transformer block for averaging every two consecutive integration approximations. As a result, BDIA-transformer can be viewed as training an ensemble of ODE solvers parameterized by a set of binary random variables, which regularizes the model and results in improved validation accuracy. Lightweight side information is required to be stored in the forward process to account for binary quantization loss to enable exact bit-level reversibility. In the inference procedure, the expectation $\mathbb{E}(\gamma)=0$ is taken to make the resulting architecture identical to transformer up to activation quantization. Our experiments in natural language generation, image classification, and language translation show that BDIA-transformers outperform their conventional counterparts significantly in terms of validation performance while also requiring considerably less training memory. Thanks to the regularizing effect of the ensemble, the BDIA-transformer is particularly suitable for fine-tuning with limited data. Source-code can be found via https://github.com/guoqiang-zhang-x/BDIA-Transformer. Guoqiang Zhang 0003, John P. Lewis, W. Bastiaan Kleijn |
ICML | 2 |
| 2024 | Directed Diffusion: Direct Control of Object Placement through Attention GuidanceabstractText-guided diffusion models such as DALLE-2, Imagen, and Stable Diffusion are able to generate an effectively endless variety of images given only a short text prompt describing the desired image content. In many cases the images are of very high quality. However, these models often struggle to compose scenes containing several key objects such as characters in specified positional relationships. The missing capability to ``direct'' the placement of characters and objects both within and across images is crucial in storytelling, as recognized in the literature on film and animation theory. In this work, we take a particularly straightforward approach to providing the needed direction. Drawing on the observation that the cross-attention maps for prompt words reflect the spatial layout of objects denoted by those words, we introduce an optimization objective that produces ``activation'' at desired positions in these cross-attention maps. The resulting approach is a step toward generalizing the applicability of text-guided diffusion models beyond single images to collections of related images, as in storybooks. Directed Diffusion provides easy high-level positional control over multiple objects, while making use of an existing pre-trained model and maintaining a coherent blend between the positioned objects and the background. Moreover, it requires only a few lines to implement. Wan-Duo Kurt Ma, Avisek Lahiri, John P. Lewis, Thomas K. Leung, W. Bastiaan Kleijn |
AAAI | 3 |
| 2024 | Exact Diffusion Inversion via Bidirectional Integration Approximation
Guoqiang Zhang 0003, John P. Lewis, W. Bastiaan Kleijn |
ECCV (57) | 2 |
| 2024 | TrailBlazer: Trajectory Control for Diffusion-Based Video Generation
Wan-Duo Kurt Ma, John P. Lewis, W. Bastiaan Kleijn |
SIGGRAPH Asia | 2 |
| 2022 | NewsStories: Illustrating Articles with Visual Summaries
Reuben Tan, Bryan A. Plummer, Kate Saenko, John P. Lewis, Avneesh Sud, Thomas K. Leung |
ECCV (36) | 4 |
| 2021 | LipSync3D: Data-Efficient Learning of Personalized 3D Talking Faces From Video Using Pose and Lighting NormalizationabstractIn this paper, we present a video-based learning framework for animating personalized 3D talking faces from audio. We introduce two training-time data normalizations that significantly improve data sample efficiency. First, we isolate and represent faces in a normalized space that decouples 3D geometry, head pose, and texture. This decomposes the prediction problem into regressions over the 3D face shape and the corresponding 2D texture atlas. Second, we leverage facial symmetry and approximate albedo constancy of skin to isolate and remove spatio-temporal lighting variations. Together, these normalizations allow simple networks to generate high fidelity lip-sync videos under novel ambient illumination while training with just a single speaker-specific video. Further, to stabilize temporal dynamics, we introduce an auto-regressive approach that conditions the model on its previous visual state. Human ratings and objective metrics demonstrate that our method outperforms contemporary state-of-the-art audio-driven video reenactment benchmarks in terms of realism, lip-sync and visual quality scores. We illustrate several applications enabled by our framework. Avisek Lahiri, Vivek Kwatra, Christian Früh, John P. Lewis, Christoph Bregler |
CVPR | 4 |
| 2020 | The HSIC Bottleneck: Deep Learning without Back-PropagationabstractWe introduce the HSIC (Hilbert-Schmidt independence criterion) bottleneck for training deep neural networks. The HSIC bottleneck is an alternative to the conventional cross-entropy loss and backpropagation that has a number of distinct advantages. It mitigates exploding and vanishing gradients, resulting in the ability to learn very deep networks without skip connections. There is no requirement for symmetric feedback or update locking. We find that the HSIC bottleneck provides performance on MNIST/FashionMNIST/CIFAR10 classification comparable to backpropagation with a cross-entropy target, even when the system is not encouraged to make the output resemble the classification labels. Appending a single layer trained with SGD (without backpropagation) to reformat the information further improves performance. Kurt Wan-Duo Ma, John P. Lewis, W. Bastiaan Kleijn |
AAAI | 2 |
| 2020 | NASA Neural Articulated Shape Approximation
Boyang Deng, John P. Lewis, Timothy Jeruzalski, Gerard Pons-Moll, Geoffrey E. Hinton, Mohammad Norouzi 0002, Andrea Tagliasacchi |
ECCV (7) | 2 |
| 2019 | High-quality object-space dynamic ambient occlusion for characters using Bi-level regressionabstractThe widely used ambient occlusion (AO) technique provides an approximation of some global illumination effects and is efficient enough for use in real-time applications. Because it relies on computing the visibility from each point on a surface, AO computation is expensive for dynamically deforming objects, such as characters in particular. In this paper, we describe an algorithm for producing high-quality dynamically changing AO for characters. Our fundamental idea is to factorize the AO computation into a coarse-scale component in which visibility is determined by approximating spheres, and a fine-scale component that leverages a skinning-like algorithm for efficiency, with both components trained in a regression against ground-truth AO values. The resulting algorithm accommodates interactions with external objects and generalizes without requiring carefully constructed training data. Extensive comparisons illustrate the capabilities and advantages of our algorithm. Binh Huy Le, Henrik Halen, Carlos Gonzalez-Ochoa, John P. Lewis |
I3D | 4 |
| 2019 | Optimal and interactive keyframe selection for motion captureabstractMotion capture is increasingly used in games and movies, but often requires editing before it can be used, for many reasons. The motion may need to be adjusted to correctly interact with virtual objects or to fix problems that result from mapping the motion to a character of a different size or, beyond such technical requirements, directors can request stylistic changes. Unfortunately, editing is laborious because of the low-level representation of the data. While existing motion editing methods accomplish modest changes, larger edits can require the artist to “re-animate” the motion by manually selecting a subset of the frames as keyframes. In this paper, we automatically find sets of frames to serve as keyframes for editing the motion. We formulate the problem of selecting an optimal set of keyframes as a shortest-path problem, and solve it efficiently using dynamic programming. We create a new simplified animation by interpolating the found keyframes using a naive curve fitting technique. Our algorithm can simplify motion capture to around 10% of the original number of frames while retaining most of its detail. By simplifying animation with our algorithm, we realize a new approach to motion editing and stylization founded on the time-tested keyframe interface. We present results that show our algorithm outperforms both research algorithms and a leading commercial tool. Richard Roberts 0004, John P. Lewis, Ken Anjyo, Jaewoo Seo, Yeongho Seol |
Comput. Vis. Media | 2 |
| 2019 | Direct delta mush skinning and variantsabstractA significant fraction of the world's population have experienced virtual characters through games and movies, and the possibility of online VR social experiences may greatly extend this audience. At present, the skin deformation for interactive and real-time characters is typically computed using geometric skinning methods. These methods are efficient and simple to implement, but obtaining quality results requires considerable manual "rigging" effort involving trial-and-error weight painting, the addition of virtual helper bones, etc. The recently introduced Delta Mush algorithm largely solves this rig authoring problem, but its iterative computational approach has prevented direct adoption in real-time engines. This paper introduces Direct Delta Mush, a new algorithm that simultaneously improves on the efficiency and control of Delta Mush while generalizing previous algorithms. Specifically, we derive a direct rather than iterative algorithm that has the same ballpark computational form as some previous geometric weight blending algorithms. Straightforward variants of the algorithm are then proposed to further optimize computational and storage cost with insignificant quality losses. These variants are equivalent to special cases of several previous skinning algorithms. Our algorithm simultaneously satisfies the goals of reasonable efficiency, quality, and ease of authoring. Further, its explicit decomposition of rotational and translational effects allows independent control over bending versus twisting deformation, as well as a skin sliding effect. Binh Huy Le, John P. Lewis |
ACM Trans. Graph. | 2 |
| 2018 | Low-Rank Matrix Completion to Reconstruct Incomplete Rendering ImagesabstractPath tracing provides photo-realistic rendering in many applications but intermediate previsualization often suffers from distracting noise. Since the fundamental underlying problem is insufficient samples, we exploit the coherence of the visual signal to reconstruct missing samples, using a low-rank matrix completion framework. We present novel methods to construct low rank matrices for incomplete images including missing pixel, missing sub-pixel, and multi-frame scenarios. A convolutional neural network provides fast pre-completion for initialising missing values, and subsequent weighted nuclear norm minimisation (WNNM) with a parameter adjustment strategy (PAWNNM) efficiently recovers missing values even in high frequency details. The result shows better visual quality than recent methods including compressed sensing based reconstruction. John P. Lewis, Taehyun Rhee |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2017 | The Shattered Gradients Problem: If resnets are the answer, then what is the question?abstractA long-standing obstacle to progress in deep learning is the problem of vanishing and exploding gradients. Although, the problem has largely been overcome via carefully constructed initializations and batch normalization, architectures incorporating skip-connections such as highway and resnets perform much better than standard feedforward architectures despite well-chosen initialization and batch normalization. In this paper, we identify the shattered gradients problem. Specifically, we show that the correlation between gradients in standard feedforward networks decays exponentially with depth resulting in gradients that resemble white noise whereas, in contrast, the gradients in architectures with skip-connections are far more resistant to shattering, decaying sublinearly. Detailed empirical evidence is presented in support of the analysis, on both fully-connected networks and convnets. Finally, we present a new “looks linear” (LL) initialization that prevents shattering, with preliminary experiments showing the new initialization allows to train very deep networks without the addition of skip-connections. David Balduzzi, Marcus Frean, Lennox Leary, John P. Lewis, Kurt Wan-Duo Ma, Brian McWilliams |
ICML | 4 |
| 2017 | Sparse Rig Parameter Optimization for Character AnimationabstractWe propose a novel motion retargeting method that efficiently estimates artist-friendly rig space parameters. Inspired by the workflow typically observed in keyframe animation, our approach transfers a source motion into a production friendly character rig by optimizing the rig space parameters while balancing the considerations of fidelity to the source motion and the ease of subsequent editing. We propose the use of an intermediate object to transfer both the skeletal motion and the mesh deformation. The target rig-space parameters are then optimized to minimize the error between the motion of an intermediate object and the target character. The optimization uses a set of artist defined weights to modulate the effect of the different rig space parameters over time. Sparsity inducing regularizers and keyframe extraction streamline any additional editing processes. The results obtained with different types of character rigs demonstrate the versatility of our method and its effectiveness in simplifying any necessary manual editing within the production pipeline. Jaewon Song, Roger Blanco Ribera, Kyungmin Cho, Mi You, John P. Lewis, Byungkuk Choi, Jun-yong Noh |
Comput. Graph. Forum | 5 |
| 2017 | Facial retargeting with automatic range of motion alignmentabstractWhile facial capturing focuses on accurate reconstruction of an actor's performance, facial animation retargeting has the goal to transfer the animation to another character, such that the semantic meaning of the animation remains. Because of the popularity of blendshape animation, this effectively means to compute suitable blendshape weights for the given target character. Current methods either require manually created examples of matching expressions of actor and target character, or are limited to characters with similar facial proportions (i.e., realistic models). In contrast, our approach can automatically retarget facial animations from a real actor to stylized characters. We formulate the problem of transferring the blendshapes of a facial rig to an actor as a special case of manifold alignment, by exploring the similarities of the motion spaces defined by the blendshapes and by an expressive training sequence of the actor. In addition, we incorporate a simple, yet elegant facial prior based on discrete differential properties to guarantee smooth mesh deformation. Our method requires only sparse correspondences between characters and is thus suitable for retargeting marker-less and marker-based motion capture as well as animation transfer between virtual characters. Roger Blanco Ribera, Eduard Zell, John P. Lewis, Jun-yong Noh, Mario Botsch |
ACM Trans. Graph. | 3 |
| 2016 | SketchiMo: sketch-based motion editing for articulated charactersabstractWe present SketchiMo, a novel approach for the expressive editing of articulated character motion. SketchiMo solves for the motion given a set of projective constraints that relate the sketch inputs to the unknown 3 D poses. We introduce the concept of sketch space, a contextual geometric representation of sketch targets---motion properties that are editable via sketch input---that enhances, right on the viewport, different aspects of the motion. The combination of the proposed sketch targets and space allows for seamless editing of a wide range of properties, from simple joint trajectories to local parent-child spatiotemporal relationships and more abstract properties such as coordinated motions. This is made possible by interpreting the user's input through a new sketch-based optimization engine in a uniform way. In addition, our view-dependent sketch space also serves the purpose of disambiguating the user inputs by visualizing their range of effect and transparently defining the necessary constraints to set the temporal boundaries for the optimization. Byungkuk Choi, Roger Blanco Ribera, John P. Lewis, Yeongho Seol, Seokpyo Hong, Haegwang Eom, Sunjin Jung, Jun-yong Noh |
ACM Trans. Graph. | 3 |
| 2012 | Spacetime expression cloning for blendshapesabstractThe goal of a practical facial animation retargeting system is to reproduce the character of a source animation on a target face while providing room for additional creative control by the animator. This article presents a novel spacetime facial animation retargeting method for blendshape face models. Our approach starts from the basic principle that the source and target movements should be similar. By interpreting movement as the derivative of position with time, and adding suitable boundary conditions, we formulate the retargeting problem as a Poisson equation. Specified (e.g., neutral) expressions at the beginning and end of the animation as well as any user-specified constraints in the middle of the animation serve as boundary conditions. In addition, a model-specific prior is constructed to represent the plausible expression space of the target face during retargeting. A Bayesian formulation is then employed to produce target animation that is consistent with the source movements while satisfying the prior constraints. Since the preservation of temporal derivatives is the primary goal of the optimization, the retargeted motion preserves the rhythm and character of the source movement and is free of temporal jitter. More importantly, our approach provides spacetime editing for the popular blendshape representation of facial models, exhibiting smooth and controlled propagation of user edits across surrounding frames. Yeongho Seol, John P. Lewis, Jaewoo Seo, Byungkuk Choi, Ken Anjyo, Jun-yong Noh |
ACM Trans. Graph. | 2 |
| 2012 | Weighted pose space editing for facial animation
Yeongho Seol, Jaewoo Seo, Paul Hyunjin Kim, John P. Lewis, Jun-yong Noh |
Vis. Comput. | 4 |
| 2011 | A Database and Evaluation Methodology for Optical FlowabstractThe quantitative evaluation of optical flow algorithms by Barron et al. (1994) led to significant advances in performance. The challenges for optical flow algorithms today go beyond the datasets and evaluation methods proposed in that paper. Instead, they center on problems associated with complex natural scenes, including nonrigid motion, real sensor noise, and motion discontinuities. We propose a new set of benchmarks and evaluation methods for the next generation of optical flow algorithms. To that end, we contribute four types of data to test different aspects of optical flow algorithms: (1) sequences with nonrigid motion where the ground-truth flow is determined by tracking hidden fluorescent texture, (2) realistic synthetic sequences, (3) high frame-rate video used to study interpolation error, and (4) modified stereo sequences of static scenes. In addition to the average angular error used by Barron et al., we compute the absolute flow endpoint error, measures for frame interpolation error, improved statistics, and results at motion discontinuities and in textureless regions. In October 2007, we published the performance of several well-known methods on a preliminary version of our data to establish the current state of the art. We also made the data freely available on the web at http://vision.middlebury.edu/flow/ . Subsequently a number of researchers have uploaded their results to our website and published papers using the data. A significant improvement in performance has already been achieved. In this paper we analyze the results obtained to date and draw a large number of conclusions from them. Simon Baker, Daniel Scharstein, John P. Lewis, Stefan Roth 0001, Michael J. Black, Richard Szeliski |
Int. J. Comput. Vis. | 3 |
| 2011 | Compression and direct manipulation of complex blendshape modelsabstractWe present a method to compress complex blendshape models and thereby enable interactive, hardware-accelerated animation of these models. Facial blendshape models in production are typically large in terms of both the resolution of the model and the number of target shapes. They are represented by a single huge blendshape matrix, whose size presents a storage burden and prevents real-time processing. To address this problem, we present a new matrix compression scheme based on a hierarchically semi-separable (HSS) representation with matrix block reordering. The compressed data are also suitable for parallel processing. An efficient GPU implementation provides very fast feedback of the resulting animation. Compared with the original data, our technique leads to a huge improvement in both storage and processing efficiency without incurring any visual artifacts. As an application, we introduce an extended version of the direct manipulation method to control a large number of facial blendshapes efficiently and intuitively. Jaewoo Seo, Geoffrey Irving, John P. Lewis, Jun-yong Noh |
ACM Trans. Graph. | 3 |
| 2011 | Artist friendly facial animation retargetingabstractThis paper presents a novel facial animation retargeting system that is carefully designed to support the animator's workflow. Observation and analysis of the animators' often preferred process of key-frame animation with blendshape models informed our research. Our retargeting system generates a similar set of blendshape weights to those that would have been produced by an animator. This is achieved by rearranging the group of blendshapes into several sequential retargeting groups and solving using a matching pursuit-like scheme inspired by a traditional key-framing approach. Meanwhile, animators typically spend a tremendous amount of time simplifying the dense weight graphs created by the retargeting. Our graph simplification technique effectively produces editable weight graphs while preserving the visual characteristics of the original retargeting. Finally, we automatically create GUI controllers to help artists perform key-framing and editing very efficiently. The set of proposed techniques greatly reduce the time and effort required by animators to achieve high quality retargeted facial animations. Yeongho Seol, Jaewoo Seo, Paul Hyunjin Kim, John P. Lewis, Jun-yong Noh |
ACM Trans. Graph. | 4 |
| 2011 | Scan-Based Volume Animation Driven by Locally Adaptive Articulated RegistrationsabstractThis paper describes a complete system to create anatomically accurate example-based volume deformation and animation of articulated body regions, starting from multiple in vivo volume scans of a specific individual. In order to solve the correspondence problem across volume scans, a template volume is registered to each sample. The wide range of pose variations is first approximated by volume blend deformation (VBD), providing proper initialization of the articulated subject in different poses. A novel registration method is presented to efficiently reduce the computation cost while avoiding strong local minima inherent in complex articulated body volume registration. The algorithm highly constrains the degrees of freedom and search space involved in the nonlinear optimization, using hierarchical volume structures and locally constrained deformation based on the biharmonic clamped spline. Our registration step establishes a correspondence across scans, allowing a data-driven deformation approach in the volume domain. The results provide an occlusion-free person-specific 3D human body model, asymptotically accurate inner tissue deformations, and realistic volume animation of articulated movements driven by standard joint control estimated from the actual skeleton. Our approach also addresses the practical issues arising in using scans from living subjects. The robustness of our algorithms is tested by their applications on the hand, probably the most complex articulated region in the body, and the knee, a frequent subject area for medical imaging due to injuries. Taehyun Rhee, John P. Lewis, Ulrich Neumann, Krishna S. Nayak |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2010 | Issues in adapting research algorithms to stereoscopic visual effectsabstractMany published machine vision algorithms are designed to be real-time and fully automatic with low computational complexity. These attributes are essential for applications such as stereo robotic vision. Motion Picture Digital Visual Effect facilities, however, have massive computation resources available and can afford human interaction to initialise algorithms and to guide them towards a good solution. On the other hand, motion pictures have significantly higher accuracy requirements and other unique challenges. Not all machine vision algorithms can readily be adapted to this environment. In this paper we outline the requirements of visual effects and indicate several challenges involved in using image processing and machine vision algorithms for stereo motion picture visual effects. Peter Hillman 0001, John P. Lewis, Sebastian Sylwan, Erik Winquist |
ICIP | 2 |
| 2010 | Stable and efficient differential inverse kinematicsabstractInverse kinematics (IK) is an essential algorithm in the computer animation of articulated figures. We propose a differential IK algorithm combining ideas from the pseudoinverse and Jacobian transpose techniques that achieves the efficiency of the former technique while avoiding its inherent instability. John P. Lewis, Nebojsa Dragosavac |
SIGGRAPH ASIA (Sketches) | 1 |
| 2010 | Scattered data interpolation and approximation for computer graphicsabstractThe goal of scattered data interpolation techniques is to construct a (typically smooth) function from a set of unorganized samples. These techniques have a wide range of applications in computer graphics. For instance they can be used to model a surface from a set of sparse samples, to reconstruct a BRDF from a set of measurements, to interpolate motion capture data, or to compute the physical properties of a fluid. This course will survey and compare scattered interpolation algorithms and describe their applications in computer graphics. Although the course is focused on applying these techniques, we will introduce some of the underlying mathematical theory and briefly mention numerical considerations. John P. Lewis, Frédéric H. Pighin, Ken Anjyo |
SIGGRAPH ASIA (Courses) | 1 |
| 2010 | A Survey of Procedural Noise FunctionsabstractAbstract Procedural noise functions are widely used in computer graphics, from off‐line rendering in movie production to interactive video games. The ability to add complex and intricate details at low memory and authoring cost is one of its main attractions. This survey is motivated by the inherent importance of noise in graphics, the widespread use of noise in industry and the fact that many recent research developments justify the need for an up‐to‐date survey. Our goal is to provide both a valuable entry point into the field of procedural noise functions, as well as a comprehensive view of the field to the informed reader. In this report, we cover procedural noise functions in all their aspects. We outline recent advances in research on this topic, discussing and comparing recent and well‐established methods. We first formally define procedural noise functions based on stochastic processes and then classify and review existing procedural noise functions. We discuss how procedural noise functions are used for modelling and how they are applied to surfaces. We then introduce analysis tools and apply them to evaluate and compare the major approaches to noise generation. We finally identify several directions for future work. Ares Lagae, Sylvain Lefebvre 0001, Robert L. Cook 0001, Tony DeRose, George Drettakis, David S. Ebert, John P. Lewis, Ken Perlin, Matthias Zwicker |
Comput. Graph. Forum | 7 |
| 2009 | Identifying salient pointsabstractThe definition of "important" or salient points on a shape is an old and fundamental problem in computer graphics. Important points, which we will term key points, can be used for representing a shape (e.g. as vertices of a polyline or as knots of a spline). Key points that correspond to perceptually salient points are useful as handles for editing. Key points are also good points to retain in simplifying a shape. John P. Lewis, Ken Anjyo |
SIGGRAPH ASIA Sketches | 1 |
| 2009 | Selecting good views of high-dimensional data using class consistencyabstractAbstract Many visualization techniques involve mapping high‐dimensional data spaces to lower‐dimensional views. Unfortunately, mapping a high‐dimensional data space into a scatterplot involves a loss of information; or, even worse, it can give a misleading picture of valuable structure in higher dimensions. In this paper, we propose class consistency as a measure of the quality of the mapping. Class consistency enforces the constraint that classes of n–D data are shown clearly in 2–D scatterplots. We propose two quantitative measures of class consistency, one based on the distance to the class's center of gravity, and another based on the entropies of the spatial distributions of classes. We performed an experiment where users choose good views, and show that class consistency has good precision and recall. We also evaluate both consistency measures over a range of data sets and show that these measures are efficient and robust. Mike Sips, Boris Neubert, John P. Lewis, Pat Hanrahan |
Comput. Graph. Forum | 3 |
| 2008 | Learning Optical Flow
Deqing Sun, Stefan Roth 0001, John P. Lewis, Michael J. Black |
ECCV (3) | 3 |
| 2007 | A Database and Evaluation Methodology for Optical FlowabstractThe quantitative evaluation of optical flow algorithms by Barron et al. led to significant advances in the performance of optical flow methods. The challenges for optical flow today go beyond the datasets and evaluation methods proposed in that paper and center on problems associated with nonrigid motion, real sensor noise, complex natural scenes, and motion discontinuities. Our goal is to establish a new set of benchmarks and evaluation methods for the next generation of optical flow algorithms. To that end, we contribute four types of data to test different aspects of optical flow algorithms: sequences with nonrigid motion where the ground-truth flow is determined by tracking hidden fluorescent texture; realistic synthetic sequences; high frame-rate video used to study interpolation error; and modified stereo sequences of static scenes. In addition to the average angular error used in Barron et al., we compute the absolute flow endpoint error, measures for frame interpolation error, improved statistics, and flow accuracy at motion boundaries and in textureless regions. We evaluate the performance of several well-known methods on this data to establish the current state of the art. Our database is freely available on the web together with scripts for scoring and publication of the results at http://vision.middlebury.edu/flow/. Simon Baker, Daniel Scharstein, John P. Lewis, Stefan Roth 0001, Michael J. Black, Richard Szeliski |
ICCV | 3 |
| 2007 | Shape Priors by Kernel Density Modeling of PCA Residual StructureabstractModern image processing techniques increasingly use prior models of the expected distribution of objects. Principal component eigen-models are often selected for shape prior modeling, but are limited in capturing only the second order moment statistics. On the other hand, kernel densities can in concept reproduce arbitrary statistics, but are problematic for high dimensional data such as shapes. An evident approach is to combine these methods, using PCA to reduce the problem dimensionality, followed by kernel density modeling of the PCA coefficients. In this paper we show that useful algorithmic and editing operations can be formulated in term of this simple approach. The operations are illustrated in the context of point distribution shape models. Particular points can be rapidly evaluated as being plausible or outliers, and a plausible shape can be completed given limited operator input in a manually guided procedure. This "PCA+KD" approach is conceptually simple, scalable (becoming increasingly accurate with additional training data), provides improved modeling power, and supports useful algorithmic queries. John P. Lewis, Iman Mostafavi, Gina E. Sosinsky, Maryann E. Martone, Ruth West |
ICIP (4) | 1 |
| 2007 | Soft-Tissue Deformation for In Vivo Volume AnimationabstractArticulated body animation with smooth skin deformation is an important topic in computer graphics. This paper presents a pipeline that extends articulated body deformation to the volume graphics domain. The pipeline consists of in-vivo volume scans, kinematic joint estimation, volumetric joint weight computation, soft-tissue volume deformation, and direct volume rendering. The result is a fully articulated body volume driven by intuitive joint control that respects rigid deformation of the bone structures and produces smooth deformations of both the skin surface and the interior soft tissue regions. Taehyun Rhee, John P. Lewis, Ulrich Neumann, Krishna S. Nayak |
PG | 2 |
| 2006 | Perceiving Visual Emotions with Speech
Zhigang Deng 0001, Jeremy N. Bailenson, John P. Lewis, Ulrich Neumann |
IVA | 3 |
| 2006 | Human hand modeling from surface anatomyabstractThe human hand is an important interface with complex shape and movement. In virtual reality and gaming applications the use of an individualized rather than generic hand representation can increase the sense of immersion and in some cases may lead to more effortless and accurate interaction with the virtual world. We present a method for constructing a person-specific model from a single canonically posed palm image of the hand without human guidance. Tensor voting is employed to extract the principal creases on the palmar surface. Joint locations are estimated using extracted features and analysis of surface anatomy. The skin geometry of a generic 3D hand model is deformed using radial basis functions guided by correspondences to the extracted surface anatomy and hand contours. The result is a 3D model of an individual's hand, with similar joint locations, contours, and skin texture. Taehyun Rhee, Ulrich Neumann, John P. Lewis |
SI3D | 3 |
| 2006 | Real-Time Weighted Pose-Space Deformation on the GPUabstractAbstract WPSD (Weighted Pose Space Deformation) is an example based skinning method for articulated body animation. The per‐vertex computation required in WPSD can be parallelized in a SIMD (Single Instruction Multiple Data) manner and implemented on a GPU. While such vertex‐parallel computation is often done on the GPU vertex processors, further parallelism can potentially be obtained by using the fragment processors. In this paper, we develop a parallel deformation method using the GPU fragment processors. Joint weights for each vertex are automatically calculated from sample poses, thereby reducing manual effort and enhancing the quality of WPSD as well as SSD (Skeletal Subspace Deformation). We show sufficient speed‐up of SSD, PSD (Pose Space Deformation) and WPSD to make them suitable for real‐time applications. Categories and Subject Descriptors (according to ACM CCS): I.3.1 [Computer Graphics]: Hardware Architecture‐Parallel processing, I.3.5 [Computer Graphics]: Computational Geometry and Object Modeling‐Curve, surface, solid and object modeling, I.3.7 [Computer Graphics]: Three‐Dimensional Graphics and Realism‐Animation. Taehyun Rhee, John P. Lewis, Ulrich Neumann |
Comput. Graph. Forum | 2 |
| 2006 | Expressive Facial Animation Synthesis by Learning Speech Coarticulation and Expression SpacesabstractSynthesizing expressive facial animation is a very challenging topic within the graphics community. In this paper, we present an expressive facial animation synthesis system enabled by automated learning from facial motion capture data. Accurate 3D motions of the markers on the face of a human subject are captured while he/she recites a predesigned corpus, with specific spoken and visual expressions. We present a novel motion capture mining technique that "learns" speech coarticulation models for diphones and triphones from the recorded data. A Phoneme-Independent Expression Eigenspace (PIEES) that encloses the dynamic expression signals is constructed by motion signal processing (phoneme-based time-warping and subtraction) and Principal Component Analysis (PCA) reduction. New expressive facial animations are synthesized as follows: First, the learned coarticulation models are concatenated to synthesize neutral visual speech according to novel speech input, then a texture-synthesis-based approach is used to generate a novel dynamic expression signal from the PIEES model, and finally the synthesized expression signal is blended with the synthesized neutral visual speech to create the final expressive facial animation. Our experiments demonstrate that the system can effectively synthesize realistic expressive facial animation. Zhigang Deng 0001, Ulrich Neumann, John P. Lewis, Tae-Yong Kim 0002, Murtaza Bulut, Shri Narayanan |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2005 | The DEFACTO System: Training Tool for Incident Commanders
Nathan Schurr, Janusz Marecki, John P. Lewis, Milind Tambe, Paul Scerri |
AAAI | 3 |
| 2005 | Synthesizing speech animation by learning compact speech co-articulation modelsabstractWhile speech animation fundamentally consists of a sequence of phonemes over time, sophisticated animation requires smooth interpolation and co-articulation effects, where the preceding and following phonemes influence the shape of a phoneme. Co-articulation has been approached in speech animation research in several ways, most often by simply smoothing the mouth geometry motion over time. Data-driven approaches tend to generate realistic speech animation, but they need to store a large facial motion database, which is not feasible for real time gaming and interactive applications on platforms such as PDAs and cell phones. In this paper we show that accurate speech co-articulation model with compact size can be learned from facial motion capture data. An initial phoneme sequence is generated automatically from text-to-speech (TTS) systems. Then, our learned co-articulation model is applied to the resulting phoneme sequence, producing natural and detailed motion. The contribution of this work is that speech co-articulation models "learned" from real human motion data can be used to generate natural-looking speech motion while simultaneously preserving the expressiveness of the animation via keyframing control. Simultaneously, this approach can be effectively applied to interactive applications due to its compact size. Zhigang Deng 0001, John P. Lewis, Ulrich Neumann |
Computer Graphics International | 2 |
| 2005 | SmartCanvas: a gesture-driven intelligent drawing desk systemabstractThis paper describes SmartCanvas, an intelligent desk system that allows a user to perform freehand drawing on a desk or similar surface with gestures. Our system requires one camera and no touch sensors. The key underlying technique is a vision-based method that distinguishes drawing gestures and transitional gestures in real time, avoiding the need for "artificial" gestures to mark the beginning and end of a drawing stroke. The method achieves an average classification accuracy of 92.17%. Pie-shaped menus and a "rotate-to-and-select" approach eliminate the need for a fixed menu display, resulting in an "invisible" interface. Zhenyao Mo, John P. Lewis, Ulrich Neumann |
IUI | 2 |
| 2005 | Reducing blendshape interference by selected motion attenuationabstractBlendshapes (linear shape interpolation models) are perhaps the most commonly employed technique in facial animation practice. A major problem in creating blendshape animation is that of blendshape interference: the adjustment of a single blendshape "slider" may degrade the effects obtained with previous slider movements, because the blendshapes have overlapping, non-orthogonal effects. Because models used in commercial practice may have 100 or more individual blendshapes, the interference problem is the subject of considerable manual effort. Modelers iteratively resculpt models to reduce interference where possible, and animators must compensate for those interference effects that remain. In this short paper we consider the blendshape interference problem from a linear algebra point of view. We find that while full orthogonality is not desirable, the goal of preserving previous adjustments to the model can be effectively approached by allowing the user to temporarily designate a set of points as representative of the previous (desired) adjustments. We then simply solve for blendshape slider values that mimic desired new movement while moving these "tagged" points as little as possible. The resulting algorithm is easy to implement and demonstrably reduces cases of blendshape interference found in existing models. John P. Lewis, Jonathan Mooser, Zhigang Deng 0001, Ulrich Neumann |
SI3D | 1 |
| 2004 | Face Inpainting with Local Linear RepresentationsabstractA number of investigators have had success using domain specific prior knowledge to produce improved superresolution images of faces ("hallucinating faces"). These efforts address the scenario where a face image is obtained from a low-resolution camera. A related but less studied problem occurs when the missing information is the result of occlusion rather than low camera resolution, as in the case when a person is wearing sunglasses. Recently Hwang and Lee [14] introduced the first algorithm for solving this reconstruction "inpainting" problem. In the current work we report results of a psychological study that provides independent evidence regarding the validity of the face reconstruction task, and we demonstrate an improved reconstruction approach using a positive, local linear representation. The positive, local mixture operates on real-world images without manual intervention in many cases, and provides demonstrably lower reconstruction error than is obtainable with a global representation. Zhenyao Mo, John P. Lewis, Ulrich Neumann |
BMVC | 2 |
| 2004 | Ripple-free local bases by designabstractIn some applications, a local or "parts based" representation is preferable to global basis functions, such as those used in Fourier and principal component analysis. In applications that require human understanding and editing of the data, it is also desirable that the basis functions be in some sense as "simple" as possible. This means, for example, that the basis functions should not have Gabor-like ripples if such ripples are not a prominent feature of the data to be represented. The paper introduces a direct local basis construction. Specifically, we show that local bases result from maximizing an appropriate redefinition of pairwise orthogonality while maintaining the ability to represent the data. The resulting basis functions are competitive with (and for some applications superior to) those obtained from existing algorithms, and the construction does not require that the basis coefficients be statistically non-Gaussian or independent, as would be the case with an independent component analysis approach. John P. Lewis, Zhenyao Mo, Ulrich Neumann |
ICASSP (3) | 1 |
| 2004 | VisualIDs: automatic distinctive icons for desktop interfacesabstractAlthough existing GUIs have a sense of space, they provide no sense of place. Numerous studies report that users misplace files and have trouble wayfinding in virtual worlds despite the fact that people have remarkable visual and spatial abilities. This issue is considered in the human-computer interface field and has been addressed with alternate display/navigation schemes. Our paper presents a fundamentally graphics based approach to this 'lost in hyperspace' problem. Specifically, we propose that spatial display of files is not sufficient to engage our visual skills; scenery (distinctive visual appearance) is needed as well. While scenery (in the form of custom icon assignments) is already possible in current operating systems, few if any users take the time to manually assign icons to all their files. As such, our proposal is to generate visually distinctive icons ("VisualIDs") automatically , while allowing the user to replace the icon if desired. The paper discusses psychological and conceptual issues relating to icons, visual memory, and the necessary relation of scenery to data. A particular icon generation algorithm is described; subjects using these icons in simulated file search and recall tasks show significantly improved performance with little effort. Although the incorporation of scenery in a graphical user interface will introduce many new (and interesting) design problems that cannot be addressed in this paper, we show that automatically created scenery is both beneficial and feasible. John P. Lewis, Ruth Rosenholtz, Nickson Fong, Ulrich Neumann |
ACM Trans. Graph. | 1 |
| 2003 | Realistic human face rendering for "The Matrix Reloaded"abstractNo abstract available. George Borshukov, John P. Lewis |
SIGGRAPH | 2 |
| 2003 | Universal capture: image-based facial animation for "The Matrix Reloaded"abstractNo abstract available. George Borshukov, Dan Piponi, Oystein Larsen, John P. Lewis, Christina Tempelaar-Lietz |
SIGGRAPH | 4 |
| 2003 | Practical eye movement model using texture synthesisabstractAs humans we are especially sensitive to the appearance of the face, and on the face, the eyes are particularly important. In fact, in attempts to animate photo-realistic CG humans, the eyes are very often what destroys the illusion [Williams 2003]. The state of art in eye movement synthesis is the Eyes Alive model [Lee et al. 2002] that develops a custom statistical model specifically for eye movement. While its results are the best to date, the model is complex and one wonders if it could be improved by using additional or different statistics. In fact the problem of generating novel animation that captures the “character” of given training data is the same problem as texture synthesis. In this sketch we describe a practical eye movement model using non-parametric texture synthesis techniques ([Efros and Leung 1999]), simulating the eye gaze motion and eye blink motion simultaneously. This approach uses the data directly and without an intervening humancrafted statistical model, yet it produces results that appear as good or better than the more complex statistical model. 2 Approach Zhigang Deng 0001, John P. Lewis, Ulrich Neumann |
SIGGRAPH | 2 |
| 2000 | Pose space deformation: a unified approach to shape interpolation and skeleton-driven deformationabstractPose space deformation generalizes and improves upon both shape interpolation and common skeleton-driven deformation techniques. This deformation approach proceeds from the observation that several types of deformation can be uniformly represented as mappings from a pose space, defined by either an underlying skeleton or a more abstract system of parameters, to displacements in the object local coordinate frames. Once this uniform representation is identified, previously disparate deformation types can be accomplished within a single unified approach. The advantages of this algorithm include improved expressive power and direct manipulation of the desired shapes yet the performance associated with traditional shape interpolation is achievable. Appropriate applications include animation of facial and body deformation for entertainment, telepresence, computer gaming, and other applications where direct sculpting of deformations is desired or where real-time synthesis of a deforming model is required. John P. Lewis, Matt Cordner, Nickson Fong |
SIGGRAPH | 1 |
| 1989 | Algorithms for solid noise synthesisabstractA solid noise is a function that defines a random value at each point in space. Solid noises have immediate and powerful applications in surface texturing, stochastic modeling, and the animation of natural phenomena.Existing solid noise synthesis algorithms are surveyed and two new algorithms are presented. The first uses Wiener interpolation to interpolate random values on a discrete lattice. The second is an efficient sparse convolution algorithm. Both algorithms are developed for model-directed synthesis, in which sampling and construction of the noise occur only at points where the noise value is required, rather than over a regularly sampled region of space. The paper attempts to present the rationale for the selection of these particular algorithms.The new algorithms have advantages of efficiency, improved control over the noise power spectrum, and the absence of artifacts. The convolution algorithm additionally allows quality to be traded for efficiency without introducing obvious deterministic effects. The algorithms are particularly suitable for applications where high-quality solid noises are required. Several sample applications in stochastic modeling and solid texturing are shown. John P. Lewis |
SIGGRAPH | 1 |
| 1987 | Automated lip-synch and speech synthesis for character animationabstractAn automated method of synchronizing facial animation to recorded speech is described. In this method, a common speech synthesis method (linear prediction) is adapted to provide simple and accurate phoneme recognition. The recognized phonemes are then associated with mouth positions to provide keyframes for computer animation of speech using a parametric model of the human face. John P. Lewis, Frederic I. Parke |
CHI | 1 |
| 1987 | Generalized Stochastic SubdivisionabstractStochastic techniques have assumed a prominent role in computer graphics because of their success in modeling a variety of complex and natural phenomena. This paper describes the basis for techniques such as stochastic subdivision in the theory of random processes and estimation theory. The popular stochastic subdivision construction is then generalized to provide control of the autocorrelation and spectral properties of the synthesized random functions. The generalized construction is suitable for generating a variety of perceptually distinct high-quality random functions, including those with non-fractal spectra and directional or oscillatory characteristics. It is argued that a spectral modeling approach provides a more powerful and somewhat more intuitive perceptual characterization of random processes than does the fractal model. Synthetic textures and terrains are presented as a means of visually evaluating the generalized subdivision technique. John P. Lewis |
ACM Trans. Graph. | 1 |
| 1984 | Texture synthesis for digital paintingabstractThe problem of digital painting is considered from a signal processing viewpoint, and is reconsidered as a problem of directed texture synthesis. It is an important characteristic of natural texture that detail may be evident at many scales, and the detail at each scale may have distinct characteristics. A “sparse convolution” procedure for generating random textures with arbitrary spectral content is described. The capability of specifying the texture spectrum (and thus the amount of detail at each scale) is an improvement over stochastic texture synthesis processes which are scalebound or which have a prescribed 1/f spectrum. This spectral texture synthesis procedure provides the basis for a digital paint system which rivals the textural sophistication of traditional artistic media. Applications in terrain synthesis and texturing computer-rendered objects are also shown. John P. Lewis |
SIGGRAPH | 1 |