EDBT 2026 Demo / reviewers in the wild / expert
Ira Kemelmacher-Shlizerman
dblp:k/IraKemelmacher · also Ira Kemelmacher
· DBLP profile ↗
50ranked-venue papers
12as first author
21since 2021 · last 2025
0009-0003-9498-584XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 42 · 11 first-author · 16 since 2021Artificial intelligence and machine learning · 35 · 10 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 6 · 5 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Perturb-and-Revise: Flexible 3D Editing with Generative TrajectoriesabstractRecent advancements in text-based diffusion models have accelerated progress in 3D reconstruction and textbased 3D editing. Although existing 3D editing methods excel at modifying color, texture, and style, they struggle with extensive geometric or appearance changes, thus limiting their applications. To this end, we propose Perturb-and-Revise, which makes possible a variety of NeRF editing. First, we perturb the NeRF parameters with random initializations to create a versatile initialization. The level of perturbation is determined automatically through analysis of the local loss landscape. Then, we revise the edited NeRF via generative trajectories. Combined with the generative process, we impose identity-preserving gradients to refine the edited NeRF. Extensive experiments demonstrate that Perturb-and-Revise facilitates flexible, effective, and consistent editing of color, appearance, and geometry in 3D. For 360° results, please visit our project page: https://susunghong.github.io/Perturb-and-Revise. Susung Hong, Johanna Suvi Karras, Ricardo Martin-Brualla, Ira Kemelmacher-Shlizerman |
CVPR | 4 |
| 2025 | Generative Inbetweening: Adapting Image-to-Video Models for Keyframe InterpolationabstractWe present a method for generating video sequences with coherent motion between a pair of input keyframes. We adapt a pretrained large-scale image-to-video diffusion model (originally trained to generate videos moving forward in time from a single input image) for keyframe interpolation, i.e., to produce a video between two input frames. We accomplish this adaptation through a lightweight fine-tuning technique that produces a version of the model that instead predicts videos moving backwards in time from a single input image. This model (along with the original forward-moving model) is subsequently used in a dual-directional diffusion sampling process that combines the overlapping model estimates starting from each of the two keyframes. Our experiments shows that our method outperforms both existing diffusion-based methods and traditional frame interpolation techniques. Brian Curless, Ira Kemelmacher-Shlizerman, Aleksander Holynski, Steven M. Seitz |
ICLR | 4 |
| 2025 | Linearly Constrained Diffusion Implicit ModelsabstractWe introduce Linearly Constrained Diffusion Implicit Models (CDIM), a fast and accurate approach to solving noisy linear inverse problems using diffusion models. Traditional diffusion-based inverse methods rely on numerous projection steps to enforce measurement consistency in addition to unconditional denoising steps. CDIM achieves a 10–50× reduction in projection steps by dynamically adjusting the number and size of projection steps to align a residual measurement energy with its theoretical distribution under the forward diffusion process. This adaptive alignment preserves measurement consistency while substantially accelerating constrained inference.
For noise-free linear inverse problems, CDIM exactly satisfies the measurement constraints with few projection steps, even when existing methods fail. We demonstrate CDIM’s effectiveness across a range of applications, including super-resolution, denoising, inpainting, deblurring, and 3D point cloud reprojection. Vivek Jayaram, Ira Kemelmacher-Shlizerman, Steven M. Seitz, John Thickstun |
NeurIPS | 2 |
| 2025 | UltraZoom: Generating Gigapixel Images from Regular PhotosabstractWe present UltraZoom, a system for generating gigapixel-resolution images of objects from casually captured inputs, such as handheld phone photos. Given a full-shot image (global, low-detail) and one or more close-ups (local, high-detail), UltraZoom upscales the full image to match the fine detail and scale of the close-up examples. To achieve this, we construct a per-instance paired dataset from the close-ups and adapt a pretrained generative model to learn object-specific low-to-high resolution mappings. At inference, we apply the model in a sliding window fashion over the full image. Constructing these pairs is non-trivial: it requires registering the close-ups within the full image for scale estimation and degradation alignment. We introduce a simple, robust method for achieving registration on arbitrary materials in casual, in-the-wild captures. Together, these components form a system that enables seamless pan and zoom across the entire object, producing consistent, photorealistic gigapixel imagery from minimal input. For full-resolution results and code, visit our project page at ultra-zoom.github.io . Vivek Jayaram, Brian Curless, Ira Kemelmacher-Shlizerman, Steven M. Seitz |
SIGGRAPH Asia | 4 |
| 2024 | Total Selfie: Generating Full-Body SelfiesabstractWe present a method to generate full-body selfies from photographs originally taken at arms length. Because self-captured photos are typically taken close up, they have lim-ited field of view and exaggerated perspective that distorts facial shapes. We instead seek to generate the photo some one else would take of you from a few feet away. Our approach takes as input four selfies of your face and body, a background image, and generates a full-body selfie in a de-sired target pose. We introduce a novel diffusion-based approach to combine all of this information into high-quality, well-composed photos of you with the desired pose and background. Bowei Chen 0003, Brian Curless, Ira Kemelmacher-Shlizerman, Steven M. Seitz |
CVPR | 3 |
| 2024 | Generative Powers of TenabstractWe present a method that uses a text-to-image model to generate consistent content across multiple image scales, enabling extreme semantic zooms into a scene, e.g. ranging from a wide-angle landscape view of a forest to a macro shot of an insect sitting on one of the tree branches. We achieve this through a joint multi-scale diffusion sampling approach that encourages consistency across different scales while preserving the integrity of each individual sampling process. Since each generated scale is guided by a different text prompt, our method enables deeper levels of zoom than traditional super-resolution methods that may struggle to create new contextual structure at vastly different scales. We compare our method qualitatively with alter-native techniques in image super-resolution and outpainting, and show that our method is most effective at generating consistent multi-scale content. Janne Kontkanen, Brian Curless, Steven M. Seitz, Ira Kemelmacher-Shlizerman, Ben Mildenhall, Pratul P. Srinivasan, Dor Verbin, Aleksander Holynski |
CVPR | 5 |
| 2024 | M&M VTO: Multi-Garment Virtual Try-On and EditingabstractWe present M &M VTO-a mix and match virtual try-on method that takes as input multiple garment images, text de-scription for garment layout and an image of a person. An example input includes: an image of a shirt, an image of a pair of pants, “rolled sleeves, shirt tucked in”, and an image of a person. The output is a visualization of how those garments (in the desired layout) would look like on the given person. Key contributions of our method are: 1) a single stage diffusion based model, with no super resolution cas-cading, that allows to mix and match multiple garments at$1024\times 512$resolution preserving and warping intricate garment details, 2) architecture design (VTO UN et Diffusion Transformer) to disentangle denoising from person specific features, allowing for a highly effective finetuning strategy for identity preservation (6MB model per individual vs 4GB achieved with, e.g., dreambooth finetuning); solving a common identity loss problem in current virtual try-on meth-ods, 3) layout control for multiple garments via text inputs finetuned over PaLI-3 [8] for virtual try-on task. Experimental results indicate that M &M VTO achieves state-of-the-art performance both qualitatively and quantitatively, as well as opens up new opportunities for virtual try-on via language-guided and multi-garment try-on. Luyang Zhu, Yingwei Li 0001, Nan Liu 0010, Ira Kemelmacher-Shlizerman |
CVPR | 6 |
| 2024 | Inverse Painting: Reconstructing The Painting Process
Bowei Chen 0003, Yifan Wang 0013, Brian Curless, Ira Kemelmacher-Shlizerman, Steven M. Seitz |
SIGGRAPH Asia | 4 |
| 2024 | Fashion-VDM: Video Diffusion Model for Virtual Try-On
Johanna Suvi Karras, Yingwei Li 0001, Nan Liu 0010, Luyang Zhu, Innfarn Yoo, Andreas Lugmayr, Ira Kemelmacher-Shlizerman |
SIGGRAPH Asia | 8 |
| 2023 | PersonNeRF : Personalized Reconstruction from Photo CollectionsabstractWe present PersonNeRF, a method that takes a collection of photos of a subject (e.g. Roger Federer) captured across multiple years with arbitrary body poses and ap-pearances, and enables rendering the subject with arbitrary novel combinations of viewpoint, body pose, and appearance. PersonNeRF builds a customized neural volumetric 3D model of the subject that is able to render an entire space spanned by camera viewpoint, body pose, and appearance. A central challenge in this task is dealing with sparse observations; a given body pose is likely only observed by a single viewpoint with a single appearance, and a given appearance is only observed under a handful of different body poses. We address this issue by recovering a canonical T-pose neural volumetric representation of the subject that allows for changing appearance across different observations, but uses a shared pose-dependent motion field across all observations. We demonstrate that this approach, along with regularization of the recovered volumetric geometry to encourage smoothness, is able to recover a model that renders compelling images from novel combinations of viewpoint, pose, and appearance from these challenging unstructured photo collections, outperforming prior work for free-viewpoint human rendering. Chung-Yi Weng, Pratul P. Srinivasan, Brian Curless, Ira Kemelmacher-Shlizerman |
CVPR | 4 |
| 2023 | TryOnDiffusion: A Tale of Two UNetsabstractGiven two images depicting a person and a garment worn by another person, our goal is to generate a visualization of how the garment might look on the input person. A key challenge is to synthesize a photorealistic detail-preserving visualization of the garment, while warping the garment to accommodate a significant body pose and shape change across the subjects. Previous methods either focus on garment detail preservation without effective pose and shape variation, or allow tryon with the desired shape and pose but lack garment details. In this paper, we propose a diffusion-based architecture that unifies two UN ets (referred to as Parallel-UNet), which allows us to preserve garment details and warp the garment for significant pose and body change in a single network. The key ideas behind Parallel-UNet include: 1) garment is warped implicitly via a cross attention mechanism, 2) garment warp and person blend happen as part of a unified process as opposed to a sequence of two separate tasks. Experimental results indicate that TryOnDiffusion achieves state-of-the-art performance both qualitatively and quantitatively. Luyang Zhu, Tyler Zhu, Fitsum Reda, Chitwan Saharia, Mohammad Norouzi 0002, Ira Kemelmacher-Shlizerman |
CVPR | 8 |
| 2023 | DreamPose: Fashion Image-to-Video Synthesis via Stable DiffusionabstractWe present DreamPose, a diffusion-based method for generating animated fashion videos from still images. Given an image and a sequence of human body poses, our method synthesizes a video containing both human and fabric motion. To achieve this, we transform a pre-trained text-to-image model (Stable Diffusion [16]) into a pose-and-image guided video synthesis model, using a novel finetuning strategy, a set of architectural changes to support the added conditioning signals, and techniques to encourage temporal consistency. We fine-tune on a collection of fashion videos from the UBC Fashion dataset [50]. We evaluate our method on a variety of clothing styles and poses, and demonstrate that our method produces state-of-the-art results on fashion video animation. Video results are available on our project page: https://grail.cs.washington.edu/projects/dreampose Johanna Suvi Karras, Aleksander Holynski, Ting-Chun Wang, Ira Kemelmacher-Shlizerman |
ICCV | 4 |
| 2023 | Animating Street ViewabstractWe present a system that automatically brings street view imagery to life by populating it with naturally behaving, animated pedestrians and vehicles. Our approach is to remove existing people and vehicles from the input image, insert moving objects with proper scale, angle, motion and appearance, plan paths and traffic behavior, as well as render the scene with plausible occlusion and shadowing effects. The system achieves these by reconstructing the still image street scene, simulating crowd behavior, and rendering with consistent lighting, visibility, occlusions, and shadows. We demonstrate results on a diverse range of street scenes including regular still images and panoramas. Mengyi Shan, Brian Curless, Ira Kemelmacher-Shlizerman, Steven M. Seitz |
SIGGRAPH Asia | 3 |
| 2023 | HRTF Estimation in the WildabstractHead Related Transfer Functions (HRTFs) play a crucial role in creating immersive spatial audio experiences. However, HRTFs differ significantly from person to person, and traditional methods for estimating personalized HRTFs are expensive, time-consuming, and require specialized equipment. We imagine a world where your personalized HRTF can be determined by capturing data through earbuds in everyday environments. In this paper, we propose a novel approach for deriving personalized HRTFs that only relies on in-the-wild binaural recordings and head tracking data. By analyzing how sounds change as the user rotates their head through different environments with different noise sources, we can accurately estimate their personalized HRTF. Our results show that our predicted HRTFs closely match ground-truth HRTFs measured in an anechoic chamber. Furthermore, listening studies demonstrate that our personalized HRTFs significantly improve sound localization and reduce front-back confusion in virtual environments. Our approach offers an efficient and accessible method for deriving personalized HRTFs and has the potential to greatly improve spatial audio experiences. Vivek Jayaram, Ira Kemelmacher-Shlizerman, Steven M. Seitz |
UIST | 2 |
| 2022 | StyleSDF: High-Resolution 3D-Consistent Image and Geometry GenerationabstractWe introduce a high resolution, 3D-consistent image and shape generation technique which we call StyleSDF. Our method is trained on single-view RGB data only, and stands on the shoulders of StyleGAN2 for image generation, while solving two main challenges in 3D-aware GANs: 1) high-resolution, view-consistent generation of the RGB images, and 2) detailed 3D shape. We achieve this by merging a SDF-based 3D representation with a style-based 2D generator. Our 3D implicit network renders low-resolution feature maps, from which the style-based network generates view-consistent, 1024×1024 images. Notably, our SDF-based 3D modeling defines detailed 3D surfaces, leading to consistent volume rendering. Our method shows higher quality results compared to state of the art in terms of visual and geometric quality. Roy Or-El, Mengyi Shan, Eli Shechtman, Jeong Joon Park, Ira Kemelmacher-Shlizerman |
CVPR | 6 |
| 2022 | HumanNeRF: Free-viewpoint Rendering of Moving People from Monocular VideoabstractWe introduce a free-viewpoint rendering method - HumanNeRF - that works on a given monocular video of a human performing complex body motions, e.g. a video from YouTube. Our method enables pausing the video at any frame and rendering the subject from arbitrary new camera viewpoints or even a full 360-degree camera path for that particular frame and body pose. This task is particularly challenging, as it requires synthesizing photorealistic details of the body, as seen from various camera angles that may not exist in the input video, as well as synthesizing fine details such as cloth folds and facial appearance. Our method optimizes for a volumetric representation of the person in a canonical T-pose, in concert with a motion field that maps the estimated canonical representation to every frame of the video via backward warps. The motion field is decomposed into skeletal rigid and non-rigid motions, produced by deep networks. We show significant performance improvements over prior work, and compelling examples of free-viewpoint renderings from monocular video of moving humans in challenging uncontrolled capture scenarios. Chung-Yi Weng, Brian Curless, Pratul P. Srinivasan, Jonathan T. Barron, Ira Kemelmacher-Shlizerman |
CVPR | 5 |
| 2022 | ClearBuds: wireless binaural earbuds for learning-based speech enhancementabstractWe present ClearBuds, the first hardware and software system that utilizes a neural network to enhance speech streamed from two wireless earbuds. Real-time speech enhancement for wireless earbuds requires high-quality sound separation and background cancellation, operating in real-time and on a mobile phone. Clear-Buds bridges state-of-the-art deep learning for blind audio source separation and in-ear mobile systems by making two key technical contributions: 1) a new wireless earbud design capable of operating as a synchronized, binaural microphone array, and 2) a lightweight dual-channel speech enhancement neural network that runs on a mobile device. Our neural network has a novel cascaded architecture that combines a time-domain conventional neural network with a spectrogram-based frequency masking neural network to reduce the artifacts in the audio output. Results show that our wireless earbuds achieve a synchronization error less than 64 μs and our network has a runtime of 21.4 ms on an accompanying mobile phone. In-the-wild evaluation with eight users in previously unseen indoor and outdoor multipath scenarios demonstrates that our neural network generalizes to learn both spatial and acoustic cues to perform noise suppression and background speech removal. In a user-study with 37 participants who spent over 15.4 hours rating 1041 audio samples collected in-the-wild, our system achieves improved mean opinion score and background noise suppression. Ishan Chatterjee, Maruchi Kim, Vivek Jayaram, Shyamnath Gollakota, Ira Kemelmacher-Shlizerman, Shwetak N. Patel, Steven M. Seitz |
MobiSys | 5 |
| 2022 | ClearBuds - wireless binaural earbuds for learning-based speech enhancementabstractWe present ClearBuds, the first end-to-end hardware and software system that utilizes a neural network to enhance speech streamed from two wireless earbuds. Real-time speech enhancement for wireless earbuds requires high-quality sound separation and background cancellation, operating in real-time and on a mobile phone. Clear-Buds bridges state-of-the-art deep learning for blind audio source separation and in-ear mobile systems by making two key technical contributions: 1) a new wireless earbud design capable of operating as a synchronized, binaural microphone array, and 2) a lightweight dual-channel speech enhancement neural network that runs on a mobile device. Our demo will allow MobiSys attendees wear our earbuds, and experience noise suppression as they talk in a noisy environment. Companion video can be accessed using the link below: Ishan Chatterjee, Maruchi Kim, Vivek Jayaram, Shyamnath Gollakota, Ira Kemelmacher-Shlizerman, Shwetak N. Patel, Steven M. Seitz |
MobiSys | 5 |
| 2021 | Real-Time High-Resolution Background MattingabstractWe introduce a real-time, high-resolution background replacement technique which operates at 30fps in 4K resolution, and 60fps for HD on a modern GPU. Our technique is based on background matting, where an additional frame of the background is captured and used in recovering the alpha matte and the foreground layer. The main challenge is to compute a high-quality alpha matte, preserving strand-level hair details, while processing high-resolution images in real-time. To achieve this goal, we employ two neural networks; a base network computes a low-resolution result which is refined by a second network operating at high-resolution on selective patches. We introduce two large-scale video and image matting datasets: VideoMatte240K and PhotoMatte13K/85. Our approach yields higher quality results compared to the previous state-of-the-art in background matting, while simultaneously yielding a dramatic boost in both speed and resolution. Shanchuan Lin, Andrey Ryabtsev, Roni Sengupta, Brian Curless, Steven M. Seitz, Ira Kemelmacher-Shlizerman |
CVPR | 6 |
| 2021 | A Light Stage on Every DeskabstractEvery time you sit in front of a TV or monitor, your face is actively illuminated by time-varying patterns of light. This paper proposes to use this time-varying illumination for synthetic relighting of your face with any new illumination condition. In doing so, we take inspiration from the light stage work of Debevec et al. [4], who first demonstrated the ability to relight people captured in a controlled lighting environment. Whereas existing light stages require expensive, room-scale spherical capture gantries and exist in only a few labs in the world, we demonstrate how to acquire useful data from a normal TV or desktop monitor. Instead of subjecting the user to uncomfortable rapidly flashing light patterns, we operate on images of the user watching a YouTube video or other standard content. We train a deep network on images plus monitor patterns of a given user and learn to predict images of that user under any target illumination (monitor pattern). Experimental evaluation shows that our method produces realistic relighting results. Video results are available at grail.cs.washington.edu/projects/Light_Stage_on_Every_Desk/. Roni Sengupta, Brian Curless, Ira Kemelmacher-Shlizerman, Steven M. Seitz |
ICCV | 3 |
| 2021 | TryOnGAN: body-aware try-on via layered interpolationabstractGiven a pair of images---target person and garment on another person---we automatically generate the target person in the given garment. Previous methods mostly focused on texture transfer via paired data training, while overlooking body shape deformations, skin color, and seamless blending of garment with the person. This work focuses on those three components, while also not requiring paired data training. We designed a pose conditioned StyleGAN2 architecture with a clothing segmentation branch that is trained on images of people wearing garments. Once trained, we propose a new layered latent space interpolation method that allows us to preserve and synthesize skin color and target body shape while transferring the garment from a different person. We demonstrate results on high resolution 512 × 512 images, and extensively compare to state of the art in try-on on both latent space generated and real images. Kathleen M. Lewis, Srivatsan Varadharajan, Ira Kemelmacher-Shlizerman |
ACM Trans. Graph. | 3 |
| 2020 | Background Matting: The World Is Your Green ScreenabstractWe propose a method for creating a matte - the per-pixel foreground color and alpha - of a person by taking photos or videos in an everyday setting with a handheld camera. Most existing matting methods require a green screen background or a manually created trimap to produce a good matte. Automatic, trimap-free methods are appearing, but are not of comparable quality. In our trimap free approach, we ask the user to take an additional photo of the background without the subject at the time of capture. This step requires a small amount of foresight but is far less timeconsuming than creating a trimap. We train a deep network with an adversarial loss to predict the matte. We first train a matting network with a supervised loss on ground truth data with synthetic composites. To bridge the domain gap to real imagery with no labeling, we train another matting network guided by the first network and by a discriminator that judges the quality of composites. We demonstrate results on a wide variety of photos and videos and show significant improvement over the state of the art. Roni Sengupta, Vivek Jayaram, Brian Curless, Steven M. Seitz, Ira Kemelmacher-Shlizerman |
CVPR | 5 |
| 2020 | Lifespan Age Transformation Synthesis
Roy Or-El, Roni Sengupta, Ohad Fried, Eli Shechtman, Ira Kemelmacher-Shlizerman |
ECCV (6) | 5 |
| 2020 | Reconstructing NBA Players
Luyang Zhu, Konstantinos Rematas, Brian Curless, Steven M. Seitz, Ira Kemelmacher-Shlizerman |
ECCV (5) | 5 |
| 2020 | The Cone of Silence: Speech Separation by LocalizationabstractGiven a multi-microphone recording of an unknown number of speakers talking concurrently, we simultaneously localize the sources and separate the individual speakers. At the core of our method is a deep network, in the waveform domain, which isolates sources within an angular region $\theta \pm w/2$, given an angle of interest $\theta$ and angular window size $w$. By exponentially decreasing $w$, we can perform a binary search to localize and separate all sources in logarithmic time. Our algorithm also allows for an arbitrary number of potentially moving speakers at test time, including more speakers than seen during training. Experiments demonstrate state of the art performance for both source separation and source localization, particularly in high levels of background noise. Teerapat Jenrungrot, Vivek Jayaram, Steven M. Seitz, Ira Kemelmacher-Shlizerman |
NeurIPS | 4 |
| 2019 | Photo Wake-Up: 3D Character Animation From a Single PhotoabstractWe present a method and application for animating a human subject from a single photo. E.g., the character can walk out, run, sit, or jump in 3D. The key contributions of this paper are: 1) an application of viewing and animating humans in single photos in 3D, 2) a novel 2D warping method to deform a posable template body model to fit the person's complex silhouette to create an animatable mesh, and 3) a method for handling partial self occlusions. We compare to state-of-the-art related methods and evaluate results with human studies. Further, we present an interactive interface that allows re-posing the person in 3D, and an augmented reality setup where the animated 3D person can emerge from the photo into the real world. We demonstrate the method on photos, posters, and art. The project page is at https://grail.cs.washington.edu/projects/wakeup/. Chung-Yi Weng, Brian Curless, Ira Kemelmacher-Shlizerman |
CVPR | 3 |
| 2018 | Soccer on Your TabletopabstractWe present a system that transforms a monocular video of a soccer game into a moving 3D reconstruction, in which the players and field can be rendered interactively with a 3D viewer or through an Augmented Reality device. At the heart of our paper is an approach to estimate the depth map of each player, using a CNN that is trained on 3D player data extracted from soccer video games. We compare with state of the art body pose and depth estimation techniques, and show results on both synthetic ground truth benchmarks, and real YouTube soccer footage. Konstantinos Rematas, Ira Kemelmacher-Shlizerman, Brian Curless, Steven M. Seitz |
CVPR | 2 |
| 2018 | Audio to Body DynamicsabstractWe present a method that gets as input an audio of violin or piano playing, and outputs a video of skeleton predictions which are further used to animate an avatar. The key idea is to create an animation of an avatar that moves their hands similarly to how a pianist or violinist would do, just from audio. Notably, it's not clear if body movement can be predicted from music at all and our aim in this work is to explore this possibility. In this paper, we present the first result that shows that natural body dynamics can be predicted. We built an LSTM network that is trained on violin and piano recital videos uploaded to the Internet. The predicted points are applied onto a rigged avatar to create the animation. Eli Shlizerman, Lucio M. Dery, Hayden Schoen, Ira Kemelmacher-Shlizerman |
CVPR | 4 |
| 2018 | Video to fully automatic 3D hair modelabstractImagine taking a selfie video with your mobile phone and getting as output a 3D model of your head (face and 3D hair strands) that can be later used in VR, AR, and any other domain. State of the art hair reconstruction methods allow either a single photo (thus compromising 3D quality) or multiple views, but they require manual user interaction (manual hair segmentation and capture of fixed camera views that span full 360°). In this paper, we describe a system that can completely automatically create a reconstruction from any video (even a selfie video), and we don't require specific views, since taking your -90°, 90°, and full back views is not feasible in a selfie capture. In the core of our system, in addition to the automatization components, hair strands are estimated and deformed in 3D (rather than 2D as in state of the art) thus enabling superior results. We provide qualitative, quantitative, and Mechanical Turk human studies that support the proposed system, and show results on a diverse variety of videos (8 different celebrity videos, 9 selfie mobile videos, spanning age, gender, hair length, type, and styling). Shu Liang, Xiufeng Huang, Xianyu Meng, Kunyao Chen, Linda G. Shapiro, Ira Kemelmacher-Shlizerman |
ACM Trans. Graph. | 6 |
| 2017 | Level Playing Field for Million Scale Face RecognitionabstractFace recognition has the perception of a solved problem, however when tested at the million-scale exhibits dramatic variation in accuracies across the different algorithms [11]. Are the algorithms very different? Is access to good/big training data their secret weapon? Where should face recognition improve? To address those questions, we created a benchmark, MF2, that requires all algorithms to be trained on same data, and tested at the million scale. MF2 is a public large-scale set with 672K identities and 4.7M photos created with the goal to level playing field for large scale face recognition. We contrast our results with findings from the other two large-scale benchmarks MegaFace Challenge and MS-Celebs-1M where groups were allowed to train on any private/public/big/small set. Some key discoveries: 1) algorithms, trained on MF2, were able to achieve state of the art and comparable results to algorithms trained on massive private sets, 2) some outperformed themselves once trained on MF2, 3) invariance to aging suffers from low accuracies as in MegaFace, identifying the need for larger age variations possibly within identities or adjustment of algorithms in future testing. Aaron Nech, Ira Kemelmacher-Shlizerman |
CVPR | 2 |
| 2017 | Synthesizing Obama: learning lip sync from audioabstractGiven audio of President Barack Obama, we synthesize a high quality video of him speaking with accurate lip sync, composited into a target video clip. Trained on many hours of his weekly address footage, a recurrent neural network learns the mapping from raw audio features to mouth shapes. Given the mouth shape at each time instant, we synthesize high quality mouth texture, and composite it with proper 3D pose matching to change what he appears to be saying in a target video to match the input audio track. Our approach produces photorealistic results. Supasorn Suwajanakorn, Steven M. Seitz, Ira Kemelmacher-Shlizerman |
ACM Trans. Graph. | 3 |
| 2016 | The MegaFace Benchmark: 1 Million Faces for Recognition at ScaleabstractRecent face recognition experiments on a major benchmark (LFW [15]) show stunning performance-a number of algorithms achieve near to perfect score, surpassing human recognition rates. In this paper, we advocate evaluations at the million scale (LFW includes only 13K photos of 5K people). To this end, we have assembled the MegaFace dataset and created the first MegaFace challenge. Our dataset includes One Million photos that capture more than 690K different individuals. The challenge evaluates performance of algorithms with increasing numbers of "distractors" (going from 10 to 1M) in the gallery set. We present both identification and verification performance, evaluate performance with respect to pose and a persons age, and compare as a function of training data size (#photos and #people). We report results of state of the art and baseline algorithms. The MegaFace dataset, baseline code, and evaluation scripts, are all publicly released for further experimentations1. Ira Kemelmacher-Shlizerman, Steven M. Seitz, Evan Brossard |
CVPR | 1 |
| 2016 | Head Reconstruction from Internet Photos
Shu Liang, Linda G. Shapiro, Ira Kemelmacher-Shlizerman |
ECCV (2) | 3 |
| 2016 | Transfiguring portraitsabstractPeople may look dramatically different by changing their hair color, hair style, when they grow older, in a different era style, or a different country or occupation. Some of those may transfigure appearance and inspire creative changes, some not, but how would we know without physically trying? We present a system that enables automatic synthesis of limitless numbers of appearances. A user inputs one or more photos (as many as they like) of his or her face, text queries an appearance of interest (just like they'd search an image search engine) and gets as output the input person in the queried appearance. Rather than fixing the number of queries or a dataset our system utilizes all the relevant and searchable images on the Internet, estimates a doppelgänger set for the inputs, and utilizes it to generate composites. We present a large number of examples on photos taken with completely unconstrained imaging conditions. Ira Kemelmacher-Shlizerman |
ACM Trans. Graph. | 1 |
| 2015 | The Meme Quiz: A Facial Expression Game Combining Human Agency and Machine Involvement
Kathleen Tuite, Ira Kemelmacher-Shlizerman |
FDG | 2 |
| 2015 | What Makes Tom Hanks Look Like Tom HanksabstractWe reconstruct a controllable model of a person from a large photo collection that captures his or her persona, i.e., physical appearance and behavior. The ability to operate on unstructured photo collections enables modeling a huge number of people, including celebrities and other well photographed people without requiring them to be scanned. Moreover, we show the ability to drive or puppeteer the captured person B using any other video of a different person A. In this scenario, B acts out the role of person A, but retains his/her own personality and character. Our system is based on a novel combination of 3D face reconstruction, tracking, alignment, and multi-texture modeling, applied to the puppeteering problem. We demonstrate convincing results on a large variety of celebrities derived from Internet imagery and video. Supasorn Suwajanakorn, Steven M. Seitz, Ira Kemelmacher-Shlizerman |
ICCV | 3 |
| 2014 | 3D Face Hallucination from a Single Depth FrameabstractWe present an algorithm that takes a single frame of a person's face from a depth camera, e.g., Kinect, and produces a high-resolution 3D mesh of the input face. We leverage a dataset of 3D face meshes of 1204 distinct individuals ranging from age 3 to 40, captured in a neutral expression. We divide the input depth frame into semantically significant regions (eyes, nose, mouth, cheeks) and search the database for the best matching shape per region. We further combine the input depth frame with the matched database shapes into a single mesh that results in a highresolution shape of the input person. Our system is fully automatic and uses only depth data for matching, making it invariant to imaging conditions. We evaluate our results using ground truth shapes, as well as compare to state-of-the-art shape estimation methods. We demonstrate the robustness of our local matching approach with high-quality reconstruction of faces that fall outside of the dataset span, e.g., faces older than 40 years old, facial expressions, and different ethnicities. Shu Liang, Ira Kemelmacher-Shlizerman, Linda G. Shapiro |
3DV | 2 |
| 2014 | Illumination-Aware Age ProgressionabstractWe present an approach that takes a single photograph of a child as input and automatically produces a series of age-progressed outputs between 1 and 80 years of age, accounting for pose, expression, and illumination. Leveraging thousands of photos of children and adults at many ages from the Internet, we first show how to compute average image subspaces that are pixel-to-pixel aligned and model variable lighting. These averages depict a prototype man and woman aging from 0 to 80, under any desired illumination, and capture the differences in shape and texture between ages. Applying these differences to a new photo yields an age progressed result. Contributions include relightable age subspaces, a novel technique for subspace-to-subspace alignment, and the most extensive evaluation of age progression techniques in the literature. Ira Kemelmacher-Shlizerman, Supasorn Suwajanakorn, Steven M. Seitz |
CVPR | 1 |
| 2014 | Total Moving Face Reconstruction
Supasorn Suwajanakorn, Ira Kemelmacher-Shlizerman, Steven M. Seitz |
ECCV (4) | 2 |
| 2013 | Internet Based Morphable ModelabstractIn this paper we present a new concept of building a morphable model directly from photos on the Internet. Morphable models have shown very impressive results more than a decade ago, and could potentially have a huge impact on all aspects of face modeling and recognition. One of the challenges, however, is to capture and register 3D laser scans of large number of people and facial expressions. Nowadays, there are enormous amounts of face photos on the Internet, large portion of which has semantic labels. We propose a framework to build a morph able model directly from photos, the framework includes dense registration of Internet photos, as well as, new single view shape reconstruction and modification algorithms. Ira Kemelmacher-Shlizerman |
ICCV | 1 |
| 2012 | Collection flowabstractComputing optical flow between any pair of Internet face photos is challenging for most current state of the art flow estimation methods due to differences in illumination, pose, and geometry. We show that flow estimation can be dramatically improved by leveraging a large photo collection of the same (or similar) object. In particular, consider the case of photos of a celebrity from Google Image Search. Any two such photos may have different facial expression, lighting and face orientation. The key idea is that instead of computing flow directly between the input pair (I, J), we compute versions of the images (I', J') in which facial expressions and pose are normalized while lighting is preserved. This is achieved by iteratively projecting each photo onto an appearance subspace formed from the full photo collection. The desired flow is obtained through concatenation of flows (I → I') o (J' → J). Our approach can be used with any two-frame optical flow algorithm, and significantly boosts the performance of the algorithm by providing invariance to lighting and shape changes. Ira Kemelmacher-Shlizerman, Steven M. Seitz |
CVPR | 1 |
| 2011 | Face reconstruction in the wildabstractWe address the problem of reconstructing 3D face models from large unstructured photo collections, e.g., obtained by Google image search or from personal photo collections in iPhoto. This problem is extremely challenging due to the high degree of variability in pose, illumination, facial expression, non-rigid changes in face shape and reflectance over time and occlusions. In light of this extreme variability, no single reconstruction can be consistent with all of the images. Instead, we define as the goal of reconstruction to recover a model that is locally consistent with the image set. I.e., each local region of the model is consistent with a large set of photos, resulting in a model that captures the dominant trends in the input data for different parts of the face. Our approach leverages multi-image shading, but unlike traditional photometric stereo approaches, allows for changes in viewpoint and shape. We optimize over pose, shape, and lighting in an iterative approach that seeks to minimize the rank of the transformed images. This approach produces high quality shape models for a wide range of celebrities from photos available on the Internet. Ira Kemelmacher-Shlizerman, Steven M. Seitz |
ICCV | 1 |
| 2011 | 3D Face Reconstruction from a Single Image Using a Single Reference Face ShapeabstractHuman faces are remarkably similar in global properties, including size, aspect ratio, and location of main features, but can vary considerably in details across individuals, gender, race, or due to facial expression. We propose a novel method for 3D shape recovery of faces that exploits the similarity of faces. Our method obtains as input a single image and uses a mere single 3D reference model of a different person's face. Classical reconstruction methods from single images, i.e., shape-from-shading, require knowledge of the reflectance properties and lighting as well as depth values for boundary conditions. Recent methods circumvent these requirements by representing input faces as combinations (of hundreds) of stored 3D models. We propose instead to use the input image as a guide to "mold" a single reference model to reach a reconstruction of the sought 3D shape. Our method assumes Lambertian reflectance and uses harmonic representations of lighting. It has been tested on images taken under controlled viewing conditions as well as on uncontrolled images downloaded from the Internet, demonstrating its accuracy and robustness under a variety of imaging conditions and overcoming significant differences in shape between the input and reference individuals including differences in facial expressions, gender, and race. Ira Kemelmacher-Shlizerman, Ronen Basri |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | Exploring photobiosabstractWe present an approach for generating face animations from large image collections of the same person. Such collections, which we call photobios , sample the appearance of a person over changes in pose, facial expression, hairstyle, age, and other variations. By optimizing the order in which images are displayed and cross-dissolving between them, we control the motion through face space and create compelling animations (e.g., render a smooth transition from frowning to smiling). Used in this context, the cross dissolve produces a very strong motion effect; a key contribution of the paper is to explain this effect and analyze its operating range. The approach operates by creating a graph with faces as nodes, and similarities as edges, and solving for walks and shortest paths on this graph. The processing pipeline involves face detection, locating fiducials (eyes/nose/mouth), solving for pose, warping to frontal views, and image comparison based on Local Binary Patterns. We demonstrate results on a variety of datasets including time-lapse photography, personal photo collections, and images of celebrities downloaded from the Internet. Our approach is the basis for the Face Movies feature in Google's Picasa. Ira Kemelmacher-Shlizerman, Eli Shechtman, Rahul Garg 0002, Steven M. Seitz |
ACM Trans. Graph. | 1 |
| 2010 | Being John Malkovich
Ira Kemelmacher-Shlizerman, Aditya Sankar, Eli Shechtman, Steven M. Seitz |
ECCV (1) | 1 |
| 2008 | 3D shape reconstruction of Mooney facesabstractTwo-tone (ldquoMooneyrdquo) images seem to arouse vivid 3D percept of faces, both familiar and unfamiliar, despite their seemingly poor content. Recent psychological and fMRI studies suggest that this percept is guided primarily by top-down procedures in which recognition precedes reconstruction. In this paper we investigate this hypothesis from a mathematical standpoint. We show that indeed, under standard shape from shading assumptions, a Mooney image can give rise to multiple different 3D reconstructions even if reconstruction is restricted to the Mooney transition curve (the boundary curve between black and white) alone. We then use top-down reconstruction methods to recover the shape of novel faces from single Mooney images exploiting prior knowledge of the structure of at least one face of a different individual. We apply these methods to thresholded images of real faces and compare the reconstruction quality relative to reconstruction from gray level images. Ira Kemelmacher-Shlizerman, Ronen Basri, Boaz Nadler |
CVPR | 1 |
| 2007 | Photometric Stereo with General, Unknown Lighting
Ronen Basri, David Jacobs 0001, Ira Kemelmacher-Shlizerman |
Int. J. Comput. Vis. | 3 |
| 2007 | A theory of locally low dimensional light transportabstractBlockwise or Clustered Principal Component Analysis (CPCA) is commonly used to achieve real-time rendering of shadows and glossy reflections with precomputed radiance transfer (PRT). The vertices or pixels are partitioned into smaller coherent regions, and light transport in each region is approximated by alocally low-dimensional subspaceusing PCA. Many earlier techniques such as surface light field and reflectance field compression use a similar paradigm. However, there has been no clear theoretical understanding of how light transport dimensionality increases with local patch size, nor of the optimal block size or number of clusters. In this paper, we develop a theory of locally low dimensional light transport, by using Szego's eigenvalue theorem to analytically derive the eigenvalues of the covariance matrix for canonical cases. We show mathematically that for symmetric patches of areaA, the number of basis functions for glossy reflections increases linearly withA, while for simple cast shadows, it often increases as √A. These results are confirmed numerically on a number of test scenes. Next, we carry out an analysis of the cost of rendering, trading off local dimensionality and the number of patches, deriving an optimal block size. Based on this analysis, we provide useful practical insights for setting parameters in CPCA and also derive a new adaptive subdivision algorithm. Moreover, we show that rendering time scales sub-linearly with the resolution of the image, allowing for interactive all-frequency relighting of 1024 x 1024 images. Dhruv Mahajan 0001, Ira Kemelmacher-Shlizerman, Ravi Ramamoorthi, Peter N. Belhumeur |
ACM Trans. Graph. | 2 |
| 2006 | Molding Face Shapes by Example
Ira Kemelmacher-Shlizerman, Ronen Basri |
ECCV (1) | 1 |
| 2005 | Indexing with Unknown Illumination and PoseabstractThe task of identifying 3D objects in 2D images is difficult due to variation in objects' appearance with changes in pose and lighting. The task is further complicated by the presence of occlusion and clutter. Shape indexing is a method for rapid association between features identified in an image and their corresponding 3D features stored in a database. Previous indexing methods ignored variations due to lighting, restricting the approach to polyhedral objects. In this paper, we further develop these methods to handle variations in both pose and lighting. We focus on rigid objects undergoing a scaled-orthographic projection and use spherical harmonics to represent lighting. The resulting integrated algorithm can recognize 3D objects from a single input image; furthermore, it recovers the pose and lighting of each familiar object in the given image. The algorithm has been tested on a database of real objects, demonstrating its performance on cluttered scenes under a variety of poses and illumination conditions. Ira Kemelmacher-Shlizerman, Ronen Basri |
CVPR (1) | 1 |