EDBT 2026 Demo / reviewers in the wild / expert
Kiriakos N. Kutulakos
dblp:64/4875 · also Kyros Kutulakos
· DBLP profile ↗
79ranked-venue papers
16as first author
12since 2021 · last 2025
0000-0002-5165-902XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 63 · 15 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 59 · 10 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Opportunistic Single-Photon Time of FlightabstractScattered light from pulsed lasers is increasingly part of our ambient illumination, as many devices rely on them for active 3D sensing. In this work, we ask: can these “ambient” light signals be detected and leveraged for passive 3D vision? We show that pulsed lasers, despite being weak and fluctuating at MHz to GHz frequencies, leave a distinctive sinc comb pattern in the temporal frequency domain of incident flux that is specific to each laser and invariant to the scene. This enables their passive detection and analysis with a free-running SPAD camera, even when they are unknown, asynchronous, out of sight, and emitting concurrently. We show how to synchronize with such lasers computationally, characterize their pulse emissions, separate their contributions, and—if many are present—localize them in 3D and recover a depth map of the camera’s field of view. We use our camera prototype to demonstrate (1) a first-of-its-kind visualization of asynchronously propagating light pulses from multiple lasers through the same scene, (2) passive estimation of a laser’s MHz-scale pulse repetition frequency with mHz precision, and (3) mm-scale 3D imaging over room-scale distances by passively harvesting photons from two or more out-of-view lasers. Sotiris Nousias, Mian Wei, Howard Xiao, Maxx Wu, Shahmeer Athar, Kevin J. Wang, Anagh Malik, David Barmherzig, David B. Lindell, Kiriakos N. Kutulakos |
CVPR | 10 |
| 2025 | Super Resolved Imaging with Adaptive Optics
Robin Swanson, Esther Y. H. Lin, Masen Lamb, Suresh Sivanandam, Kiriakos N. Kutulakos |
ICCV | 5 |
| 2025 | GStex: Per-Primitive Texturing of 2D Gaussian Splatting for Decoupled Appearance and Geometry ModelingabstractGaussian splatting has demonstrated excellent performance for view synthesis and scene reconstruction. The representation achieves photorealistic quality by optimizing the position, scale, color, and opacity of thousands to millions of 2D or 3D Gaussian primitives within a scene. However, since each Gaussian primitive encodes both appearance and geometry, these attributes are strongly coupled-thus, high-fidelity appearance modeling requires a large number of Gaussian primitives, even when the scene geometry is simple (e.g., for a textured planar surface). We propose to texture each 2D Gaussian primitive so that even a single Gaussian can be used to capture appearance details. By employing per-primitive texturing, our appearance representation is agnostic to the topology and complexity of the scene's geometry. We show that our approach, GStex, yields improved visual quality over prior work in texturing Gaussian splats. Furthermore, we demonstrate that our decoupling enables improved novel view synthesis performance compared to 2D Gaussian splatting when reducing the number of Gaussian primitives, and that GStex can be used for scene appearance editing and re-texturing. Victor Rong, Jingxiang Chen, Sherwin Bahmani, Kiriakos N. Kutulakos, David B. Lindell |
WACV | 4 |
| 2025 | Learning Lens Blur FieldsabstractOptical blur is an inherent property of any lens system and is challenging to model in modern cameras because of their complex optical elements. To tackle this challenge, we introduce a high-dimensional neural representation of blur-the lens blur field-and a practical method for acquiring it. The lens blur field is a multilayer perceptron (MLP) designed to (1) accurately capture variations of the lens 2D point spread function over image plane location, focus setting and, optionally, depth and (2) represent these variations parametrically as a single, sensor-specific function. The representation models the combined effects of defocus, diffraction, aberration, and accounts for sensor features such as pixel color filters and pixel-specific micro-lenses. To learn the real-world blur field of a given device, we formulate a generalized non-blind deconvolution problem that directly optimizes the MLP weights using a small set of focal stacks as the only input. We also provide a first-of-its-kind dataset of 5D blur fields-for smartphone cameras, camera bodies equipped with a variety of lenses, etc. Lastly, we show that acquired 5D blur fields are expressive and accurate enough to reveal, for the first time, differences in optical behavior of smartphone devices of the same make and model. Esther Y. H. Lin, Zhecheng Wang 0001, Rebecca Lin, Daniel Miau, Florian Kainz, Jiawen Chen 0001, Xuaner Cecilia Zhang, David B. Lindell, Kiriakos N. Kutulakos |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2025 | Generating the Past, Present and Future from a Motion-Blurred ImageabstractWe seek to answer the question: what can a motion-blurred image reveal about a scene's past, present, and future? Although motion blur obscures image details and degrades visual quality, it also encodes information about scene and camera motion during an exposure. Previous techniques leverage this information to estimate a sharp image from an input blurry one, or to predict a sequence of video frames showing what might have occurred at the moment of image capture. However, they rely on handcrafted priors or network architectures to resolve ambiguities in this inverse problem, and do not incorporate image and video priors on large-scale datasets. As such, existing methods struggle to reproduce complex scene dynamics and do not attempt to recover what occurred before or after an image was taken. Here, we introduce a new technique that repurposes a pre-trained video diffusion model trained on internet-scale datasets to recover videos revealing complex scene dynamics during the moment of capture and what might have occurred immediately into the past or future. Our approach is robust and versatile; it outperforms previous methods for this task, generalizes to challenging in-the-wild images, and supports downstream tasks such as recovering camera trajectories, object motion, and dynamic 3D scene structure. Code and data are available at blur2vid.github.io SaiKiran Kumar Tedla, Kelly Zhu, Trevor D. Canham, Felix Taubner, Michael S. Brown, Kiriakos N. Kutulakos, David B. Lindell |
ACM Trans. Graph. | 6 |
| 2024 | TurboSL: Dense, Accurate and Fast 3D by Neural Inverse Structured LightabstractWe show how to turn a noisy and fragile active triangulation technique-three-pattern structured light with a grayscale camera-into a fast and powerful tool for 3D capture: able to output sub-pixel accurate disparities at megapixel resolution, along with reflectance, normals, and a no-reference estimate of its own pixelwise 3D error. To achieve this, we formulate structured-light decoding as a neural inverse rendering problem. We show that despite having just three or four input images-all from the same viewpoint-this problem can be tractably solved by TurboSL, an algorithm that combines (1) a precise image formation model, (2) a signed distance field scene representation, and (3) projection Pattern sequences optimized for accuracy instead of precision. We use TurboSL to reconstruct a variety of complex scenes from images captured at up to 60 fps with a camera and a common projector. Our experiments highlight TurboSL's potential for dense and highly-accurate 3D acquisition from data captured in fractions of a second. Parsa Mirdehghan, Maxx Wu, Wenzheng Chen, David B. Lindell, Kiriakos N. Kutulakos |
CVPR | 5 |
| 2024 | Flying with Photons: Rendering Novel Views of Propagating Light
Anagh Malik, Noah Juravsky, Ryan Po, Gordon Wetzstein, Kiriakos N. Kutulakos, David B. Lindell |
ECCV (21) | 5 |
| 2024 | Coherent Optical Modems for Full-Wavefield Lidar
Parsa Mirdehghan, Brandon Buscaino, Maxx Wu, Doug Charlton, Mohammad E. Mousa-Pasandi, Kiriakos N. Kutulakos, David B. Lindell |
SIGGRAPH Asia | 6 |
| 2023 | Passive Ultra-Wideband Single-Photon ImagingabstractWe consider the problem of imaging a dynamic scene over an extreme range of timescales simultaneously—seconds to picoseconds—and doing so passively, without much light, and without any timing signals from the light source(s) emitting it. Because existing flux estimation techniques for single-photon cameras break down in this regime, we develop a flux probing theory that draws insights from stochastic calculus to enable reconstruction of a pixel’s time-varying flux from a stream of monotonically-increasing photon detection timestamps. We use this theory to (1) show that passive free-running SPAD cameras have an attainable frequency bandwidth that spans the entire DC-to-31 GHz range in low-flux conditions, (2) derive a novel Fourier-domain flux reconstruction algorithm that scans this range for frequencies with statistically-significant support in the timestamp data, and (3) ensure the algorithm’s noise model remains valid even for very low photon counts or non-negligible dead times. We show the potential of this asynchronous imaging regime by experimentally demonstrating several never-seen-before abilities: (1) imaging a scene illuminated simultaneously by sources operating at vastly different speeds without synchronization (bulbs, projectors, multiple pulsed lasers), (2) passive non-line-of-sight video acquisition, and (3) recording ultra-wideband video, which can be played back later at 30 Hz to show everyday motions—but can also be played a billion times slower to show the propagation of light itself. Mian Wei, Sotiris Nousias, Rahul Gulve, David B. Lindell, Kiriakos N. Kutulakos |
ICCV | 5 |
| 2023 | Transient Neural Radiance Fields for Lidar View Synthesis and 3D ReconstructionabstractNeural radiance fields (NeRFs) have become a ubiquitous tool for modeling scene appearance and geometry from multiview imagery. Recent work has also begun to explore how to use additional supervision from lidar or depth sensor measurements in the NeRF framework. However, previous lidar-supervised NeRFs focus on rendering conventional camera imagery and use lidar-derived point cloud data as auxiliary supervision; thus, they fail to incorporate the underlying image formation model of the lidar. Here, we propose a novel method for rendering transient NeRFs that take as input the raw, time-resolved photon count histograms measured by a single-photon lidar system, and we seek to render such histograms from novel views. Different from conventional NeRFs, the approach relies on a time-resolved version of the volume rendering equation to render the lidar measurements and capture transient light transport phenomena at picosecond timescales. We evaluate our method on a first-of-its-kind dataset of simulated and captured transient multiview scans from a prototype single-photon lidar. Overall, our work brings NeRFs to a new dimension of imaging at transient timescales, newly enabling rendering of transient imagery from novel views. Additionally, we show that our approach recovers improved geometry and conventional appearance compared to point cloud-based supervision when training on few input viewpoints. Transient NeRFs may be especially useful for applications which seek to simulate raw lidar measurements for downstream tasks in autonomous driving, robotics, and remote sensing. Anagh Malik, Parsa Mirdehghan, Sotiris Nousias, Kiriakos N. Kutulakos, David B. Lindell |
NeurIPS | 4 |
| 2023 | Discrete Search Photometric Stereo for Fast and Accurate Shape EstimationabstractWe consider the problem of estimating surface normals of a scene with spatially varying, general bidirectional reflectance distribution functions (BRDFs) observed by a static camera under varying distant illuminations. Unlike previous approaches that rely on continuous optimization of surface normals, we cast the problem as a discrete search problem over a set of finely discretized surface normals. In this setting, we show that the expensive processes can be precomputed in a scene-independent manner, resulting in accelerated inference. We discuss two variants of our discrete search photometric stereo (DSPS), one working with continuous linear combinations of BRDF bases and the other working with discrete BRDFs sampled from a BRDF space. Experiments show that DSPS has comparable accuracy to state-of-the-art exemplar-based photometric stereo methods while achieving 10-100x acceleration. Kenji Enomoto, Michael Waechter, Fumio Okura, Kiriakos N. Kutulakos, Yasuyuki Matsushita |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Computational Imaging on the Electric GridabstractNight beats with alternating current (AC) illumination. By passively sensing this beat, we reveal new scene information which includes: the type of bulbs in the scene, the phases of the electric grid up to city scale, and the light transport matrix. This information yields unmixing of reflections and semi-reflections, nocturnal high dynamic range, and scene rendering with bulbs not observed during acquisition. The latter is facilitated by a dataset of bulb response functions for a range of sources, which we collected and provide. To do all this, we built a novel coded-exposure high-dynamic-range imaging technique, specifically designed to operate on the grid's AC lighting. Mark Sheinin, Yoav Y. Schechner, Kiriakos N. Kutulakos |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | A Neural Rendering Framework for Free-Viewpoint RelightingabstractWe present a novel Relightable Neural Renderer (RNR) for simultaneous view synthesis and relighting using multi-view image inputs. Existing neural rendering (NR) does not explicitly model the physical rendering process and hence has limited capabilities on relighting. RNR instead models image formation in terms of environment lighting, object intrinsic attributes, and light transport function (LTF), each corresponding to a learnable component. In particular, the incorporation of a physically based rendering process not only enables relighting but also improves the quality of view synthesis. Comprehensive experiments on synthetic and real data show that RNR provides a practical and effective solution for conducting free-viewpoint relighting. Anpei Chen, Guli Zhang, Yu Ji 0001, Kiriakos N. Kutulakos, Jingyi Yu 0001 |
CVPR | 6 |
| 2020 | Auto-Tuning Structured Light by Optical Stochastic Gradient DescentabstractWe consider the problem of optimizing the performance of an active imaging system by automatically discovering the illuminations it should use, and the way to decode them. Our approach tackles two seemingly incompatible goals: (1) ''tuning'' the illuminations and decoding algorithm precisely to the devices at hand---to their optical transfer functions, non-linearities, spectral responses, image processing pipelines---and (2) doing so without modeling or calibrating the system; without modeling the scenes of interest; and without prior training data. The key idea is to formulate a stochastic gradient descent (SGD) optimization procedure that puts the actual system in the loop: projecting patterns, capturing images, and calculating the gradient of expected reconstruction error. We apply this idea to structured-light triangulation to ''auto-tune'' several devices---from smartphones and laser projectors to advanced computational cameras. Our experiments show that despite being model-free and automatic, optical SGD can boost system 3D accuracy substantially over state-of-the-art coding schemes. Wenzheng Chen, Parsa Mirdehghan, Sanja Fidler, Kiriakos N. Kutulakos |
CVPR | 4 |
| 2020 | Photometric Stereo via Discrete Hypothesis-and-Test SearchabstractIn this paper, we consider the problem of estimating surface normals of a scene with spatially varying, general BRDFs observed by a static camera under varying, known, distant illumination. Unlike previous approaches that are mostly based on continuous local optimization, we cast the problem as a discrete hypothesis-and-test search problem over the discretized space of surface normals. While a naive search requires a significant amount of time, we show that the expensive computation block can be precomputed in a scene-independent manner, resulting in accelerated inference for new scenes. It allows us to perform a full search over the finely discretized space of surface normals to determine the globally optimal surface normal for each scene point. We show that our method can accurately estimate surface normals of scenes with spatially varying different reflectances in a reasonable amount of time. Kenji Enomoto, Michael Waechter, Kiriakos N. Kutulakos, Yasuyuki Matsushita |
CVPR | 3 |
| 2020 | End-to-End Video Compressive Sensing Using Anderson-Accelerated Unrolled NetworksabstractCompressive imaging systems with spatial-temporal encoding can be used to capture and reconstruct fast-moving objects. The imaging quality highly depends on the choice of encoding masks and reconstruction methods. In this paper, we present a new network architecture to jointly design the encoding masks and the reconstruction method for compressive high-frame-rate imaging. Unlike previous works, the proposed method takes full advantage of denoising prior to provide a promising frame reconstruction. The network is also flexible enough to optimize full-resolution masks and efficient at reconstructing frames. To this end, we develop a new dense network architecture that embeds Anderson acceleration, known from numerical optimization, directly into the neural network architecture. Our experiments show the optimized masks and the dense accelerated network respectively achieve 1.5 dB and 1 dB improvements in PSNR without adding training parameters. The proposed method outperforms other state-of-the-art methods both in simulations and on real hardware. In addition, we set up a coded two-bucket camera for compressive high-frame-rate imaging, which is robust to imaging noise and provides promising results when recovering nearly 1,000 frames per second. Miao Qi, Rahul Gulve, Mian Wei, Roman Genov, Kiriakos N. Kutulakos, Wolfgang Heidrich |
ICCP | 6 |
| 2020 | Learned feature embeddings for non-line-of-sight imaging and recognitionabstractObjects obscured by occluders are considered lost in the images acquired by conventional camera systems, prohibiting both visualization and understanding of such hidden objects. Non-line-of-sight methods (NLOS) aim at recovering information about hidden scenes, which could help make medical imaging less invasive, improve the safety of autonomous vehicles, and potentially enable capturing unprecedented high-definition RGB-D data sets that include geometry beyond the directly visible parts. Recent NLOS methods have demonstrated scene recovery from time-resolved pulse-illuminated measurements encoding occluded objects as faint indirect reflections. Unfortunately, these systems are fundamentally limited by the quartic intensity fall-off for diffuse scenes. With laser illumination limited by eye-safety limits, recovery algorithms must tackle this challenge by incorporating scene priors. However, existing NLOS reconstruction algorithms do not facilitate learning scene priors. Even if they did, datasets that allow for such supervision do not exist, and successful encoder-decoder networks and generative adversarial networks fail for real-world NLOS data. In this work, we close this gap by learning hidden scene feature representations tailored to both reconstruction and recognition tasks such as classification or object detection, while still relying on physical models at the feature level. We overcome the lack of real training data with a generalizable architecture that can be trained in simulation. We learn the differentiable scene representation jointly with the reconstruction task using a differentiable transient renderer in the objective, and demonstrate that it generalizes to unseen classes and unseen real-world scenes , unlike existing encoder-decoder architectures and generative adversarial networks. The proposed method allows for end-to-end training for different NLOS tasks , such as image reconstruction, classification, and object detection, while being memory-efficient and running at real-time rates. We demonstrate hidden view synthesis, RGB-D reconstruction, classification, and object detection in the hidden scene in an end-to-end fashion. Wenzheng Chen, Fangyin Wei, Kiriakos N. Kutulakos, Szymon Rusinkiewicz, Felix Heide |
ACM Trans. Graph. | 3 |
| 2019 | A Theory of Fermat Paths for Non-Line-Of-Sight Shape ReconstructionabstractWe present a novel theory of Fermat paths of light between a known visible scene and an unknown object not in the line of sight of a transient camera. These light paths either obey specular reflection or are reflected by the object's boundary, and hence encode the shape of the hidden object. We prove that Fermat paths correspond to discontinuities in the transient measurements. We then derive a novel constraint that relates the spatial derivatives of the path lengths at these discontinuities to the surface normal. Based on this theory, we present an algorithm, called Fermat Flow, to estimate the shape of the non-line-of-sight object. Our method allows, for the first time, accurate shape recovery of complex objects, ranging from diffuse to specular, that are hidden around the corner as well as hidden behind a diffuser. Finally, our approach is agnostic to the particular technology used for transient imaging. As such, we demonstrate mm-scale shape recovery from pico-second scale transients using a SPAD and ultrafast laser, as well as micron-scale reconstruction from femto-second scale transients using interferometry. We believe our work is a significant advance over the state-of-the-art in non-line-of-sight imaging. Shumian Xin, Sotiris Nousias, Kiriakos N. Kutulakos, Aswin C. Sankaranarayanan, Srinivasa G. Narasimhan, Ioannis Gkioulekas |
CVPR | 3 |
| 2018 | Optimal Structured Light à La CarteabstractWe consider the problem of automatically generating sequences of structured-light patterns for active stereo triangulation of a static scene. Unlike existing approaches that use predetermined patterns and reconstruction algorithms tied to them, we generate patterns on the fly in response to generic specifications: number of patterns, projector-camera arrangement, workspace constraints, spatial frequency content, etc. Our pattern sequences are specifically optimized to minimize the expected rate of correspondence errors under those specifications for an unknown scene, and are coupled to a sequence-independent algorithm for perpixel disparity estimation. To achieve this, we derive an objective function that is easy to optimize and follows from first principles within a maximum-likelihood framework. By minimizing it, we demonstrate automatic discovery of pattern sequences, in under three minutes on a laptop, that can outperform state-of-the-art triangulation techniques. Parsa Mirdehghan, Wenzheng Chen, Kiriakos N. Kutulakos |
CVPR | 3 |
| 2018 | Coded Two-Bucket Cameras for Computer Vision
Mian Wei, Navid Sarhangnejad, Zhengfan Xia, Nikita Gusev, Nikola Katic, Roman Genov, Kiriakos N. Kutulakos |
ECCV (3) | 7 |
| 2018 | Rolling shutter imaging on the electric gridabstractFlicker of AC-powered lights is useful for probing the electric grid and unmixing reflected contributions of different sources. Flicker has been sensed in great detail with a specially-designed camera tethered to an AC outlet. We argue that even an untethered smartphone can achieve the same task. We exploit the inter-row exposure delay of the ubiquitous rolling-shutter sensor. When pixel exposure time is kept short, this delay creates a spatiotemporal wave pattern that encodes (1) the precise capture time relative to the AC, (2) the response function of individual bulbs, and (3) the AC phase that powers them. To sense point sources, we induce the spatiotemporal wave pattern by placing a star filter or a paper diffuser in front of the camera's lens. We demonstrate several new capabilities, including: high-rate acquisition of bulb response functions from one smartphone photo; recognition of bulb type and phase from one or two images; and rendering of live flicker video, as if it came from a high speed global-shutter camera. Mark Sheinin, Yoav Y. Schechner, Kiriakos N. Kutulakos |
ICCP | 3 |
| 2017 | Computational Imaging on the Electric GridabstractNight beats with alternating current (AC) illumination. By passively sensing this beat, we reveal new scene information which includes: the type of bulbs in the scene, the phases of the electric grid up to city scale, and the light transport matrix. This information yields unmixing of reflections and semi-reflections, nocturnal high dynamic range, and scene rendering with bulbs not observed during acquisition. The latter is facilitated by a database of bulb response functions for a range of sources, which we collected and provide. To do all this, we built a novel coded-exposure high-dynamic-range imaging technique, specifically designed to operate on the grids AC lighting. Mark Sheinin, Yoav Y. Schechner, Kiriakos N. Kutulakos |
CVPR | 3 |
| 2017 | Depth from Defocus in the WildabstractWe consider the problem of two-frame depth from defocus in conditions unsuitable for existing methods yet typical of everyday photography: a non-stationary scene, a handheld cellphone camera, a small aperture, and sparse scene texture. The key idea of our approach is to combine local estimation of depth and flow in very small patches with a global analysis of image content-3D surfaces, deformations, figure-ground relations, textures. To enable local estimation we (1) derive novel defocus-equalization filters that induce brightness constancy across frames and (2) impose a tight upper bound on defocus blur-just three pixels in radius-by appropriately refocusing the camera for the second input frame. For global analysis we use a novel splinebased scene representation that can propagate depth and flow across large irregularly-shaped regions. Our experiments show that this combination preserves sharp boundaries and yields good depth and flow maps in the face of significant noise, non-rigidity, and data sparsity. Huixuan Tang, Scott Cohen, Brian L. Price, Stephen Schiller, Kiriakos N. Kutulakos |
CVPR | 5 |
| 2017 | The Geometry of First-Returning Photons for Non-Line-of-Sight ImagingabstractNon-line-of-sight (NLOS) imaging utilizes the full 5D light transient measurements to reconstruct scenes beyond the cameras field of view. Mathematically, this requires solving an elliptical tomography problem that unmixes the shape and albedo from spatially-multiplexed measurements of the NLOS scene. In this paper, we propose a new approach for NLOS imaging by studying the properties of first-returning photons from three-bounce light paths. We show that the times of flight of first-returning photons are dependent only on the geometry of the NLOS scene and each observation is almost always generated from a single NLOS scene point. Exploiting these properties, we derive a space carving algorithm for NLOS scenes. In addition, by assuming local planarity, we derive an algorithm to localize NLOS scene points in 3D and estimate their surface normals. Our methods do not require either the full transient measurements or solving the hard elliptical tomography problem. We demonstrate the effectiveness of our methods through simulations as well as real data captured from a SPAD sensor. Chia-Yin Tsai, Kiriakos N. Kutulakos, Srinivasa G. Narasimhan, Aswin C. Sankaranarayanan |
CVPR | 2 |
| 2017 | Epipolar time-of-flight imagingabstractConsumer time-of-flight depth cameras like Kinect and PMD are cheap, compact and produce video-rate depth maps in short-range applications. In this paper we apply energy-efficient epipolar imaging to the ToF domain to significantly expand the versatility of these sensors: we demonstrate live 3D imaging at over 15 m range outdoors in bright sunlight; robustness to global transport effects such as specular and diffuse inter-reflections---the first live demonstration for this ToF technology; interference-free 3D imaging in the presence of many ToF sensors, even when they are all operating at the same optical wavelength and modulation frequency; and blur-free, distortion-free 3D video in the presence of severe camera shake. We believe these achievements can make such cheap ToF devices broadly applicable in consumer and robotics domains. Supreeth Achar, Joseph R. Bartels, William Whittaker, Kiriakos N. Kutulakos, Srinivasa G. Narasimhan |
ACM Trans. Graph. | 4 |
| 2016 | 3D Shape and Indirect Appearance by Structured Light TransportabstractWe consider the problem of deliberately manipulating the direct and indirect light flowing through a time-varying, general scene in order to simplify its visual analysis. Our approach rests on a crucial link between stereo geometry and light transport: while direct light always obeys the epipolar geometry of a projector-camera pair, indirect light overwhelmingly does not. We show that it is possible to turn this observation into an imaging method that analyzes light transport in real time in the optical domain, prior to acquisition. This yields three key abilities that we demonstrate in an experimental camera prototype: (1) producing a live indirect-only video stream for any scene, regardless of geometric or photometric complexity; (2) capturing images that make existing structured-light shape recovery algorithms robust to indirect transport; and (3) turning them into one-shot methods for dynamic 3D shape capture. Matthew O'Toole, John Mather, Kiriakos N. Kutulakos |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | Defocus deblurring and superresolution for time-of-flight depth camerasabstractContinuous-wave time-of-flight (ToF) cameras show great promise as low-cost depth image sensors in mobile applications. However, they also suffer from several challenges, including limited illumination intensity, which mandates the use of large numerical aperture lenses, and thus results in a shallow depth of field, making it difficult to capture scenes with large variations in depth. Another shortcoming is the limited spatial resolution of currently available ToF sensors. In this paper we analyze the image formation model for blurred ToF images. By directly working with raw sensor measurements but regularizing the recovered depth and amplitude images, we are able to simultaneously deblur and super-resolve the output of ToF cameras. Our method outperforms existing methods on both synthetic and real datasets. In the future our algorithm should extend easily to cameras that do not follow the cosine model of continuous-wave sensors, as well as to multi-frequency or multi-phase imaging employed in more recent ToF cameras. Lei Xiao 0014, Felix Heide, Matthew O'Toole, Andreas Kolb 0001, Matthias B. Hullin, Kiriakos N. Kutulakos, Wolfgang Heidrich |
CVPR | 6 |
| 2015 | High Resolution Photography with an RGB-Infrared CameraabstractA convenient solution to RGB-Infrared photography is to extend the basic RGB mosaic with a fourth filter type with high transmittance in the near-infrared band. Unfortunately, applying conventional demosaicing algorithms to RGB-IR sensors is not possible for two reasons. First, the RGB and near-infrared image are differently focused due to different refractive indices of each band. Second, manufacturing constraints introduce crosstalk between RGB and IR channels. In this paper we propose a novel image formation model for RGB-IR cameras that can be easily calibrated, and propose an efficient algorithm that jointly addresses three restoration problems--channel deblurring, channel separation and pixel demosaicing--using quadratic image regularizers. We also extend our algorithm to handle more general regularizers and pixel saturation. Experiments show that our method produces sharp, full-resolution images of pure RGB color and IR. Huixuan Tang, Xiaopeng Zhang 0001, Shaojie Zhuo, Kiriakos N. Kutulakos, Liang Shen 0007 |
ICCP | 5 |
| 2015 | Homogeneous codes for energy-efficient illumination and imagingabstractProgrammable coding of light between a source and a sensor has led to several important results in computational illumination, imaging and display. Little is known, however, about how to utilize energy most effectively, especially for applications in live imaging. In this paper, we derive a novel framework to maximize energy efficiency by "homogeneous matrix factorization" that respects the physical constraints of many coding mechanisms (DMDs/LCDs, lasers, etc. ). We demonstrate energy-efficient imaging using two prototypes based on DMD and laser illumination. For our DMD-based prototype, we use fast local optimization to derive codes that yield brighter images with fewer artifacts in many transport probing tasks. Our second prototype uses a novel combination of a low-power laser projector and a rolling shutter camera. We use this prototype to demonstrate never-seen-before capabilities such as (1) capturing live structured-light video of very bright scenes---even a light bulb that has been turned on; (2) capturing epipolar-only and indirect-only live video with optimal energy efficiency; (3) using a low-power projector to reconstruct 3D objects in challenging conditions such as strong indirect light, strong ambient light, and smoke; and (4) recording live video from a projector's---rather than the camera's---point of view. Matthew O'Toole, Supreeth Achar, Srinivasa G. Narasimhan, Kiriakos N. Kutulakos |
ACM Trans. Graph. | 4 |
| 2014 | 3D Shape and Indirect Appearance by Structured Light TransportabstractWe consider the problem of deliberately manipulating the direct and indirect light flowing through a time-varying, fully-general scene in order to simplify its visual analysis. Our approach rests on a crucial link between stereo geometry and light transport: while direct light always obeys the epipolar geometry of a projector-camera pair, indirect light overwhelmingly does not. We show that it is possible to turn this observation into an imaging method that analyzes light transport in real time in the optical domain, prior to acquisition. This yields three key abilities that we demonstrate in an experimental camera prototype: (1) producing a live indirect-only video stream for any scene, regardless of geometric or photometric complexity, (2) capturing images that make existing structured-light shape recovery algorithms robust to indirect transport, and (3) turning them into one-shot methods for dynamic 3D shape capture. Matthew O'Toole, John Mather, Kiriakos N. Kutulakos |
CVPR | 3 |
| 2014 | Temporal frequency probing for 5D transient analysis of global light transportabstractWe analyze light propagation in an unknown scene using projectors and cameras that operate at transient timescales. In this new photography regime, the projector emits a spatio-temporal 3D signal and the camera receives a transformed version of it, determined by the set of all light transport paths through the scene and the time delays they induce. The underlying 3D-to-3D transformation encodes scene geometry and global transport in great detail, but individual transport components ( e.g ., direct reflections, inter-reflections, caustics, etc .) are coupled nontrivially in both space and time. To overcome this complexity, we observe that transient light transport is always separable in the temporal frequency domain . This makes it possible to analyze transient transport one temporal frequency at a time by trivially adapting techniques from conventional projector-to-camera transport. We use this idea in a prototype that offers three never-seen-before abilities: (1) acquiring time-of-flight depth images that are robust to general indirect transport, such as interreflections and caustics; (2) distinguishing between direct views of objects and their mirror reflection; and (3) using a photonic mixer device to capture sharp, evolving wavefronts of "light-in-flight". Matthew O'Toole, Felix Heide, Lei Xiao 0014, Matthias B. Hullin, Wolfgang Heidrich, Kiriakos N. Kutulakos |
ACM Trans. Graph. | 6 |
| 2013 | What does an aberrated photo tell us about the lens and the scene?abstractWe investigate the feasibility of recovering lens properties, scene appearance and depth from a single photo containing optical aberrations and defocus blur. Starting from the ray intersection function of a rotationally-symmetric compound lens and the theory of Seidel aberrations, we obtain three basic results. First, we derive a model for the lens PSF that (1) accounts for defocus and primary Seidel aberrations and (2) describes how light rays are bent by the lens. Second, we show that the problem of inferring depth and aberration coefficients from the blur kernel of just one pixel has three degrees of freedom in general. As such it cannot be solved unambiguously. Third, we show that these degrees of freedom can be eliminated by inferring scaled aberration coefficients and depth from the blur kernel at multiple pixels in a single photo (at least three). These theoretical results suggest that single-photo aberration estimation and depth recovery may indeed be possible, given the recent progress on blur kernel estimation and blind deconvolution. Huixuan Tang, Kiriakos N. Kutulakos |
ICCP | 2 |
| 2012 | Utilizing Optical Aberrations for Extended-Depth-of-Field Panoramas
Huixuan Tang, Kiriakos N. Kutulakos |
ACCV (4) | 2 |
| 2012 | Frequency Analysis of Transient Light Transport with Applications in Bare Sensor Imaging
Di Wu 0006, Gordon Wetzstein, Christopher Barsi, Thomas Willwacher, Matthew O'Toole, Nikhil Naik 0003, Qionghai Dai, Kiriakos N. Kutulakos, Ramesh Raskar |
ECCV (1) | 8 |
| 2012 | Primal-dual coding to probe light transportabstractWe present primal-dual coding , a photography technique that enables direct fine-grain control over which light paths contribute to a photo. We achieve this by projecting a sequence of patterns onto the scene while the sensor is exposed to light. At the same time, a second sequence of patterns, derived from the first and applied in lockstep, modulates the light received at individual sensor pixels. We show that photography in this regime is equivalent to a matrix probing operation in which the elements of the scene's transport matrix are individually re-scaled and then mapped to the photo. This makes it possible to directly acquire photos in which specific light transport paths have been blocked, attenuated or enhanced. We show captured photos for several scenes with challenging light transport effects, including specular inter-reflections, caustics, diffuse inter-reflections and volumetric scattering. A key feature of primal-dual coding is that it operates almost exclusively in the optical domain: our results consist of directly-acquired, unprocessed RAW photos or differences between them. Matthew O'Toole, Ramesh Raskar, Kiriakos N. Kutulakos |
ACM Trans. Graph. | 3 |
| 2011 | Light-Efficient PhotographyabstractIn this paper, we consider the problem of imaging a scene with a given depth of field at a given exposure level in the shortest amount of time possible. We show that by 1) collecting a sequence of photos and 2) controlling the aperture, focus, and exposure time of each photo individually, we can span the given depth of field in less total time than it takes to expose a single narrower-aperture photo. Using this as a starting point, we obtain two key results. First, for lenses with continuously variable apertures, we derive a closed-form solution for the globally optimal capture sequence, i.e., that collects light from the specified depth of field in the most efficient way possible. Second, for lenses with discrete apertures, we derive an integer programming problem whose solution is the optimal sequence. Our results are applicable to off-the-shelf cameras and typical photography conditions, and advocate the use of dense, wide-aperture photo sequences as a light-efficient alternative to single-shot, narrow-aperture photography. Samuel W. Hasinoff, Kiriakos N. Kutulakos |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Dynamic Refraction StereoabstractIn this paper we consider the problem of reconstructing the 3D position and surface normal of points on an unknown, arbitrarily-shaped refractive surface. We show that two viewpoints are sufficient to solve this problem in the general case, even if the refractive index is unknown. The key requirements are 1) knowledge of a function that maps each point on the two image planes to a known 3D point that refracts to it, and 2) light is refracted only once. We apply this result to the problem of reconstructing the time-varying surface of a liquid from patterns placed below it. To do this, we introduce a novel "stereo matching" criterion called refractive disparity, appropriate for refractive scenes, and develop an optimization-based algorithm for individually reconstructing the position and normal of each point projecting to a pixel in the input views. Results on reconstructing a variety of complex, deforming liquid surfaces suggest that our technique can yield detailed reconstructions that capture the dynamic behavior of free-flowing liquids. Nigel J. W. Morris, Kiriakos N. Kutulakos |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Non-rigid structure from locally-rigid motionabstractWe introduce locally-rigid motion, a general framework for solving the M-point, N-view structure-from-motion problem for unknown bodies deforming under orthography. The key idea is to first solve many local 3-point, N-view rigid problems independently, providing a “soup” of specific, plausibly rigid, 3D triangles. The main advantage here is that the extraction of 3D triangles requires only very weak assumptions: (1) deformations can be locally approximated by near-rigid motion of three points (i.e., stretching not dominant) and (2) local motions involve some generic rotation in depth. Triangles from this soup are then grouped into bodies, and their depth flips and instantaneous relative depths are determined. Results on several sequences, both our own and from related work, suggest these conditions apply in diverse settings - including very challenging ones (e.g., multiple deforming bodies). Our starting point is a novel linear solution to 3-point structure from motion, a problem for which no general algorithms currently exist. Allan Douglas Jepson, Kiriakos N. Kutulakos |
CVPR | 3 |
| 2010 | Transparent and Specular Object ReconstructionabstractAbstract This state of the art report covers reconstruction methods for transparent and specular objects or phenomena. While the 3D acquisition of opaque surfaces with Lambertian reflectance is a well‐studied problem, transparent, refractive, specular and potentially dynamic scenes pose challenging problems for acquisition systems. This report reviews and categorizes the literature in this field. Despite tremendous interest in object digitization, the acquisition of digital models of transparent or specular objects is far from being a solved problem. On the other hand, real‐world data is in high demand for applications such as object modelling, preservation of historic artefacts and as input to data‐driven modelling techniques. With this report we aim at providing a reference for and an introduction to the field of transparent and specular object reconstruction. We describe acquisition approaches for different classes of objects. Transparent objects/phenomena that do not change the straight ray geometry can be found foremost in natural phenomena. Refraction effects are usually small and can be considered negligible for these objects. Phenomena as diverse as fire, smoke, and interstellar nebulae can be modelled using a straight ray model of image formation. Refractive and specular surfaces on the other hand change the straight rays into usually piecewise linear ray paths, adding additional complexity to the reconstruction problem. Translucent objects exhibit significant sub‐surface scattering effects rendering traditional acquisition approaches unstable. Different classes of techniques have been developed to deal with these problems and good reconstruction results can be achieved with current state‐of‐the‐art techniques. However, the approaches are still specialized and targeted at very specific object classes. We classify the existing literature and hope to provide an entry point to this exiting field. Ivo Ihrke, Kiriakos N. Kutulakos, Hendrik P. A. Lensch, Marcus A. Magnor, Wolfgang Heidrich |
Comput. Graph. Forum | 2 |
| 2010 | Linear Sequence-to-Sequence AlignmentabstractIn this paper, we consider the problem of estimating the spatiotemporal alignment between N unsynchronized video sequences of the same dynamic 3D scene, captured from distinct viewpoints. Unlike most existing methods, which work for N = 2 and rely on a computationally intensive search in the space of temporal alignments, we present a novel approach that reduces the problem for general N to the robust estimation of a single line in IR(N). This line captures all temporal relations between the sequences and can be computed without any prior knowledge of these relations. Considering that the spatial alignment is captured by the parameters of fundamental matrices, an iterative algorithm is used to refine simultaneously the parameters representing the temporal and spatial relations between the sequences. Experimental results with real-world and synthetic sequences show that our method can accurately align the videos even when they have large misalignments (e.g., hundreds of frames), when the problem is seemingly ambiguous (e.g., scenes with roughly periodic motion), and when accurate manual alignment is difficult (e.g., due to slow-moving objects). Flávio L. C. Pádua, Rodrigo L. Carceroni, Geraldo A. M. R. Santos, Kiriakos N. Kutulakos |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2010 | Optical computing for fast light transport analysisabstractWe present a general framework for analyzing the transport matrix of a real-world scene at full resolution, without capturing many photos. The key idea is to use projectors and cameras to directly acquire eigenvectors and the Krylov subspace of the unknown transport matrix. To do this, we implement Krylov subspace methods partially in optics, by treating the scene as a "black box subroutine" that enables optical computation of arbitrary matrix-vector products. We describe two methods--- optical Arnoldi to acquire a low-rank approximation of the transport matrix for relighting; and optical GMRES to invert light transport. Our experiments suggest that good quality relighting and transport inversion are possible from a few dozen low-dynamic range photos, even for scenes with complex shadows, caustics, and other challenging lighting effects. Matthew O'Toole, Kiriakos N. Kutulakos |
ACM Trans. Graph. | 2 |
| 2009 | Time-constrained photographyabstractCapturing multiple photos at different focus settings is a powerful approach for reducing optical blur, but how many photos should we capture within a fixed time budget? We develop a framework to analyze optimal capture strategies balancing the tradeoff between defocus and sensor noise, incorporating uncertainty in resolving scene depth. We derive analytic formulas for restoration error and use Monte Carlo integration over depth to derive optimal capture strategies for different camera designs, under a wide range of photographic scenarios. We also derive a new upper bound on how well spatial frequencies can be preserved over the depth of field. Our results show that by capturing the optimal number of photos, a standard camera can achieve performance at the level of more complex computational cameras, in all but the most demanding of cases. We also show that computational cameras, although specifically designed to improve one-shot performance, generally benefit from capturing multiple photos as well. Samuel W. Hasinoff, Kiriakos N. Kutulakos, Frédo Durand, William T. Freeman |
ICCV | 2 |
| 2009 | Confocal Stereo
Samuel W. Hasinoff, Kiriakos N. Kutulakos |
Int. J. Comput. Vis. | 2 |
| 2009 | TurboPixels: Fast Superpixels Using Geometric FlowsabstractWe describe a geometric-flow-based algorithm for computing a dense oversegmentation of an image, often referred to as superpixels. It produces segments that, on one hand, respect local image boundaries, while, on the other hand, limiting undersegmentation through a compactness constraint. It is very fast, with complexity that is approximately linear in image size, and can be applied to megapixel sized images with high superpixel densities in a matter of minutes. We show qualitative demonstrations of high-quality results on several complex images. The Berkeley database is used to quantitatively compare its performance to a number of oversegmentation algorithms, showing that it yields less undersegmentation than algorithms that lack a compactness constraint while offering a significant speedup over N-cuts, which does enforce compactness. Alex Levinshtein, Adrian Stere, Kiriakos N. Kutulakos, David J. Fleet, Sven J. Dickinson, Kaleem Siddiqi |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2008 | Semidefinite Programming Heuristics for Surface Reconstruction Ambiguities
Ady Ecker, Allan Douglas Jepson, Kiriakos N. Kutulakos |
ECCV (1) | 3 |
| 2008 | Light-Efficient Photography
Samuel W. Hasinoff, Kiriakos N. Kutulakos |
ECCV (4) | 2 |
| 2008 | A Theory of Refractive and Specular 3D Shape by Light-Path Triangulation
Kiriakos N. Kutulakos, Eron Steger |
Int. J. Comput. Vis. | 1 |
| 2007 | Shape from Planar Curves: A Linear Escape from FlatlandabstractWe revisit the problem of recovering 3D shape from the projection of planar curves on a surface. This problem is strongly motivated by perception studies. Applications include single-view modeling and fully uncalibrated structured light. When the curves intersect, the problem leads to a linear system for which a direct least-squares method is sensitive to noise. We derive a more stable solution and show examples where the same method produces plausible surfaces from the projection of parallel (non-intersecting) planar cross sections. Ady Ecker, Kiriakos N. Kutulakos, Allan Douglas Jepson |
CVPR | 2 |
| 2007 | A Layer-Based Restoration Framework for Variable-Aperture PhotographyabstractWe present variable-aperture photography, a new method for analyzing sets of images captured with different aperture settings, with all other camera parameters fixed. We show that by casting the problem in an image restoration framework, we can simultaneously account for defocus, high dynamic range exposure (HDR), and noise, all of which are confounded according to aperture. Our formulation is based on a layered decomposition of the scene that models occlusion effects in detail. Recovering such a scene representation allows us to adjust the camera parameters in post-capture, to achieve changes in focus setting or depth-of-field—with all results available in HDR. Our method is designed to work with very few input images: we demonstrate results from real sequences obtained using the three-image "aperture bracketing" mode found on consumer digital SLR cameras. Samuel W. Hasinoff, Kiriakos N. Kutulakos |
ICCV | 2 |
| 2007 | Reconstructing the Surface of Inhomogeneous Transparent Scenes by Scatter-Trace PhotographyabstractWe present a new method for reconstructing the exterior surface of a complex transparent scene with inhomogeneous interior (e.g., multiple interfaces, reflective or painted interiors, etc). Our approach involves capturing images of the scene from one or more viewpoints while moving a proximal light source to a 2D or 3D set of positions. This gives a 2D (or 3D) dataset per pixel, called the scatter trace. The key idea of our approach is that even though light transport within a transparent scene's interior can be exceedingly complex, the scatter trace of each pixel has a highly-constrained geometry that (1) reveals the contribution of direct surface reflection, and (2) leads to a simple "scatter- trace stereo" algorithm for computing the local geometry of the exterior surface (depth and surface normals). We present 3D reconstruction results for a variety of scenes that exhibit complex light transport phenomena. Nigel J. W. Morris, Kiriakos N. Kutulakos |
ICCV | 2 |
| 2007 | Photo-Consistent Reconstruction of Semitransparent Scenes by Density-Sheet DecompositionabstractThis paper considers the problem of reconstructing visually realistic 3D models of dynamic semitransparent scenes, such as fire, from a very small set of simultaneous views (even two). We show that this problem is equivalent to a severely underconstrained computerized tomography problem, for which traditional methods break down. Our approach is based on the observation that every pair of photographs of a semitransparent scene defines a unique density field, called a Density Sheet, that 1) concentrates all its density on one connected, semitransparent surface, 2) reproduces the two photos exactly, and 3) is the most spatially compact density field that does so. From this observation, we reduce reconstruction to the convex combination of sheet-like density fields, each of which is derived from the Density Sheet of two input views. We have applied this method specifically to the problem of reconstructing 3D models of fire. Experimental results suggest that this method enables high-quality view synthesis without overfitting artifacts. Samuel W. Hasinoff, Kiriakos N. Kutulakos |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Confocal Stereo
Samuel W. Hasinoff, Kiriakos N. Kutulakos |
ECCV (1) | 2 |
| 2005 | A Theory of Refractive and Specular 3D Shape by Light-Path TriangulationabstractWe investigate the feasibility of reconstructing an arbitrarily-shaped specular scene (refractive or mirror-like) from one or more viewpoints. By reducing shape recovery to the problem of reconstructing individual 3D light paths that cross the image plane, we obtain three key results. First, we show how to compute the depth map of a specular scene from a single viewpoint, when the scene redirects incoming light just once. Second, for scenes where incoming light undergoes two refractions or reflections, we show that three viewpoints are sufficient to enable reconstruction in the general case. Third, we show that it is impossible to reconstruct individual light paths when light is redirected more than twice. Our analysis assumes that, for every point on the image plane, we know at least one 3D point on its light path. This leads to reconstruction algorithms that rely on an "environment matting" procedure to establish pixel-to-point correspondences along a light path. Preliminary results for a variety of scenes (mirror, glass, etc) are also presented. Kiriakos N. Kutulakos, Eron Steger |
ICCV | 1 |
| 2005 | Dynamic Refraction StereoabstractIn this paper we consider the problem of reconstructing the 3D position and surface normal of points on an unknown, arbitrarily-shaped refractive surface. We show that two viewpoints are sufficient to solve this problem in the general case, even if the refractive index is unknown. The key requirements are: (1) knowledge of a function that maps each point on the two image planes to a known 3D point that refracts to it; and (2) light is refracted only once. We apply this result to the problem of reconstructing the time-varying surface of a liquid from patterns placed below it. To do this, we introduce a novel stereo matching criterion called refractive disparity, appropriate for refractive scenes, and develop an optimization-based algorithm for individually reconstructing the position and normal of each point projecting to a pixel in the input views. Results on reconstructing a variety of complex, deforming liquid surfaces suggest that our technique can yield detailed reconstructions that capture the dynamic behavior of free-flowing liquids Nigel J. W. Morris, Kiriakos N. Kutulakos |
ICCV | 2 |
| 2005 | A Theory of Inverse Light TransportabstractIn this paper we consider the problem of computing and removing interreflections in photographs of real scenes. Towards this end, we introduce the problem of inverse light transport - given a photograph of an unknown scene, decompose it into a sum of n-bounce images, where each image records the contribution of light that bounces exactly n times before reaching the camera. We prove the existence of a set of interreflection cancelation operators that enable computing each n-bounce image by multiplying the photograph by a matrix. This matrix is derived from a set of "impulse images" obtained by probing the scene with a narrow beam of light. The operators work under unknown and arbitrary illumination, and exist for scenes that have arbitrary spatially-varying BRDFs. We derive a closed-form expression for these operators in the Lambertian case and present experiments with textured and untextured Lambertian scenes that confirm our theory's predictions. Steven M. Seitz, Yasuyuki Matsushita, Kiriakos N. Kutulakos |
ICCV | 3 |
| 2004 | Linear Sequence-to-Sequence Alignment
Rodrigo L. Carceroni, Flávio L. C. Pádua, Geraldo A. M. R. Santos, Kiriakos N. Kutulakos |
CVPR (1) | 4 |
| 2003 | Photo-Consistent 3D Fire by Flame-Sheet DecompositionabstractThis paper considers the problem of reconstructing visually realistic 3D models of fire from a very small set of simultaneous views (even two). By modeling fire as a semitransparent 3D density field, we show that fire reconstruction is equivalent to a severely under-constrained computerized tomography problem, for which traditional methods break down. Our approach is based on the observation that every pair of photographs of a semitransparent scene defines a unique density field, called a Flame Sheet, that (1) concentrates all its density on one connected, semitransparent surface, (2) reproduces the two photos exactly, and (3) is the most spatially-coherent density field that does so. From this observation, we reduce fire reconstruction to the convex combination of sheet-like density fields, each of which is derived from the Flame Sheet of two input photos. Experimental results suggest that this method enables high-quality view extrapolation without over-fitting artifacts. Samuel W. Hasinoff, Kiriakos N. Kutulakos |
ICCV | 2 |
| 2002 | A Probabilistic Theory of Occupancy and Emptiness
Rahul Bhotika, David J. Fleet, Kiriakos N. Kutulakos |
ECCV (3) | 3 |
| 2002 | Multi-View Scene Capture by Surfel Sampling: From Video Streams to Non-Rigid 3D Motion, Shape and Reflectance
Rodrigo L. Carceroni, Kiriakos N. Kutulakos |
Int. J. Comput. Vis. | 2 |
| 2002 | Guest Editorial
Kiriakos N. Kutulakos, Amnon Shashua |
Int. J. Comput. Vis. | 1 |
| 2002 | Plenoptic Image Editing
Steven M. Seitz, Kiriakos N. Kutulakos |
Int. J. Comput. Vis. | 2 |
| 2001 | Multi-View Scene Capture by Surfel Sampling: From Video Streams to Non-Rigid 3D MotionShape & ReflectanceabstractIn this paper we study the problem of recovering the 3D shape, reflectance, and non-rigid motion of a dynamic 3D scene. Because these properties are completely unknown, our approach uses multiple views to build a piecewise continuous geometric and radiometric representation of the scene's trace in space-time. Basic primitive of this representation is the dynamic surfel, which (1) encodes the instantaneous local shape, reflectance, and motion of a small region in the scene, and (2) enables accurate prediction of the region's dynamic appearance under known illumination conditions. We show that complete surfel-based reconstructions can be created by repeatedly applying an algorithm called surfel sampling that combines sampling and parameter estimation to fit a single surfel to a small, bounded region of space-time. Experimental results with the Phong reflectance model and complex real scenes (clothing, skin, shiny objects) illustrate our method's ability to explain pixels and pixel variations in terms of their physical causes-shape, reflectance, motion, illumination, and visibility. Rodrigo L. Carceroni, Kiriakos N. Kutulakos |
ICCV | 2 |
| 2000 | Approximate N-View Stereo
Kiriakos N. Kutulakos |
ECCV (1) | 1 |
| 2000 | A Theory of Shape by Space Carving
Kiriakos N. Kutulakos, Steven M. Seitz |
Int. J. Comput. Vis. | 1 |
| 1999 | Toward Recovering Shape and Motion of 3D Curves from Multi-View Image SequencesabstractWe introduce a framework for recovering the 3D shape and motion of unknown, arbitrarily-moving curves from two or more image sequences acquired simultaneously from distinct points in space. We use this framework to (1) identify ambiguities in the multi-view recovery of (rigid or nonrigid) 3D motion for arbitrary curves, and (2) identify a novel spatio-temporal constraint that couples the problems of 3D shape and 3D motion recovery in the multi-view case. We show that this constraint leads to a simple hypothesize-and-test algorithm for estimating 3D curve shape and motion simultaneously. Experiments performed with synthetic data suggest that, in addition to recovering 3D curve motion, our approach yields shape estimates of higher accuracy than those obtained when stereo analysis alone is applied to a multi-view sequence. Rodrigo L. Carceroni, Kiriakos N. Kutulakos |
CVPR | 2 |
| 1999 | Multi-View 3D Shape and Motion Recovery on the Spatio-Temporal Curve ManifoldabstractIn this paper we consider the problem of recovering the 3D motion and shape of an arbitrarily-moving, arbitrarily-shaped curve from multiple synchronized video streams acquired from distinct and known points in space. By studying the 3D motion and shape constraints provided by the input video streams, we show that (1) shape and motion recovery is equivalent to the problem of recovering the differential properties of the spatio-temporal curve manifold that describes the curve's trace in space-time, and (2) a local analytical description of this manifold can be computed directly from the spatio-temporal volumes defined by the input video streams. Our experimental results suggest that this manifold-based approach to joint shape and motion estimation yields shape estimates of higher accuracy that those obtained from stereo alone, allows accurate recovery of 3D curve motion, and provides significant robustness against image noise and camera calibration errors. Rodrigo L. Carceroni, Kiriakos N. Kutulakos |
ICCV | 2 |
| 1999 | A Theory of Shape by Space CarvingabstractIn this paper we consider the problem of computing the 3D shape of an unknown, arbitrarily-shaped scene from multiple photographs taken at known but arbitrarily-distributed viewpoints. By studying the equivalence class of all 3D shapes that reproduce the input photographs, we prove the existence of a special member of this class, the photo hull, that (1) can be computed directly from photographs of the scene, and (2) subsumes all other members of this class. We then give a provably-correct algorithm called Space Carving, for computing this shape and present experimental results on complex real-world scenes. The approach is designed to (1) build photorealistic shapes that accurately model scene appearance from a wide range of viewpoints, and (2) account for the complex interactions between occlusion, parallax, shading, and their effects on arbitrary views of a 3D scene. Kiriakos N. Kutulakos, Steven M. Seitz |
ICCV | 1 |
| 1998 | Plenoptic Image EditingabstractThis paper presents a new class of interactive image editing operations designed to maintain consistency between multiple images of a physical 3D scene. The distinguishing feature of these operations is that edits to any one image propagate automatically to all other images as if the (unknown) 3D scene had itself been modified. The modified scene can then be viewed interactively from any other camera viewpoint and under different scene illuminations. The approach is useful first as a power-assist that enables a user to quickly modify many images by editing just a few, and second as a means for constructing and editing image-based scene representations by manipulating a set of photographs. The approach works by extending operations like image painting, scissoring, and morphing so that they alter a scene's generalized plenoptic function in a physically-consistent way, thereby affecting scene appearance from all viewpoints simultaneously. A key element in realizing these operations is a new volumetric decomposition technique for reconstructing an scene's plenoptic function from an incomplete set of camera viewpoints. Steven M. Seitz, Kiriakos N. Kutulakos |
ICCV | 2 |
| 1998 | Calibration-Free Augmented RealityabstractCamera calibration and the acquisition of Euclidean 3D measurements have so far been considered necessary requirements for overlaying three-dimensional graphical objects with live video. We describe a new approach to video-based augmented reality that avoids both requirements: it does not use any metric information about the calibration parameters of the camera or the 3D locations and dimensions of the environment's objects. The only requirement is the ability to track across frames at least four fiducial points that are specified by the user during system initialization and whose world coordinates are unknown. Our approach is based on the following observation: given a set of four or more noncoplanar 3D points, the projection of all points in the set can be computed as a linear combination of the projections of just four of the points. We exploit this observation by: tracking regions and color fiducial points at frame rate; and representing virtual objects in a non-Euclidean, affine frame of reference that allows their projection to be computed as a linear combination of the projection of the fiducial points. Experimental results on two augmented reality systems, one monitor-based and one head-mounted, demonstrate that the approach is readily implementable, imposes minimal computational and hardware requirements, and generates real-time and accurate video overlays even when the camera parameters vary dynamically. Kiriakos N. Kutulakos, James R. Vallino |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 1997 | Shape from the Light Field BoundaryabstractRay-based representations of shape have received little attention in computer vision. In this paper we show that the problem of recovering shape from silhouettes becomes considerably simplified if it is formulated as a reconstruction problem in the space of oriented rays that intersect the object. The method can be used with both calibrated and uncalibrated cameras, does not rely on point correspondences to compute shape, and does not impose restrictions on object topology or smoothness. Kiriakos N. Kutulakos |
CVPR | 1 |
| 1995 | Affine Surface Reconstruction by Purposive Viewpoint ControlabstractWe present an approach for building an affine representation of an unknown curved object viewed under orthographic projection from images of its occluding contour. It is based on the observation that the projection of a point on a curved, featureless surface can be computed along a special viewing direction that does not belong to the point's tangent plane. We show that by circumnavigating the object on the tangent plane of selected surface points, we can (1) compute two orthogonal projections of every point projecting to the occluding contour during this motion, and (2) compute the affine coordinates of these points. Our approach demonstrates that affine shape of curved objects can be computed directly, i.e., without Euclidean calibration or image velocity and acceleration measurements.> Kiriakos N. Kutulakos |
ICCV | 1 |
| 1995 | Global Surface Reconstruction by Purposive Control of Observer Motion
Kiriakos N. Kutulakos, Charles R. Dyer |
Artif. Intell. | 1 |
| 1994 | Occluding contour detection using affine invariants and purposive viewpoint controlabstractWe present an approach for identifying the occluding contour and determining its sidedness using an active (i.e., moving) observer. It is based on the non-stationarity property of the visible rim: When the observer's viewpoint is changed, the visible rim is a collection of curves that "slide," rigidly or non-rigidly over the surface. We show that the absenter can deterministically choose three views on the tangent plane of selected surface points to distinguish such curves from stationary surface curves (i.e., surface markings). Our approach demonstrates that the occluding contour can be identified directly, i.e., without first computing surface shape (distance and curvature).> Kiriakos N. Kutulakos, Charles R. Dyer |
CVPR | 1 |
| 1994 | Global surface reconstruction by purposive control of observer motionabstractWhat real-time, qualitative viewpoint-control behaviors are important for performing global visual exploration tasks such as searching for specific surface markings, building a global model of an arbitrary object, or recognizing an object? In this paper we consider the task of purposefully controlling the motion of an active, monocular observer in order to recover a global description of a smooth, arbitrarily-shaped object using the occluding contour. By studying the epipolar parameterization, we develop two basic behaviors that allow reconstruction of a patch around any point in a reconstructible surface region. These behaviors rely only on information extracted directly from images (e.g., tangents to the occluding contour), and are simple enough to be executed in real time. We then show how global surface reconstruction can be provably achieved by (1) integrating these behaviors to iteratively "grow" the reconstructed regions, and (2) obeying four simple rules.> Kiriakos N. Kutulakos, Charles R. Dyer |
CVPR | 1 |
| 1994 | Provable Strategies for Vision-Guided Exploration in Three DimensionsabstractAn approach is presented for exploring an unknown, arbitrary surface in three-dimensional (3D) space by a mobile robot. The main contributions are (1) an analysis of the capabilities a robot must possess and the trade-offs involved in the design of an exploration strategy, and (2) two provably-correct exploration strategies that exploit these trade-offs and use visual sensors (e.g., cameras and range sensors) to plan the robot's motion. No such analysis existed previously for the case of a robot moving freely in 3D space. The approach exploits the notion of the occlusion boundary, i.e., the points separating the visible from the occluded parts of an object. The occlusion boundary is a collection of curves that "slide" over the surface when the robot's position is continuously controlled, inducing the visibility of surface points over which they slide. The paths generated by our strategies force the occlusion boundary to slide over the entire surface. The strategies provide a basis for integrating motion planning and visual sensing under a common computational framework.> Kiriakos N. Kutulakos, Charles R. Dyer, Vladimir J. Lumelsky |
ICRA | 1 |
| 1994 | Recovering shape by purposive viewpoint adjustment
Kiriakos N. Kutulakos, Charles R. Dyer |
Int. J. Comput. Vis. | 1 |
| 1993 | Toward global surface reconstruction by purposive viewpoint adjustmentabstractThe following problem is considered: how should an observer change viewpoint in order to generate a dense image sequence of an arbitrary smooth surface so that it can be incrementally reconstructed using the occluding contour and the epipolar parameterization? A collection of qualitative behaviors is presented that, when integrated appropriately, purposefully control viewpoint based on the appearance of the surface in order to provably solve this problem.> Kiriakos N. Kutulakos, Charles R. Dyer |
CVPR | 1 |
| 1992 | Recovering shape by purposive viewpoint adjustmentabstractAn approach for recovering surface shape from the occluding contour using an active (i.e., moving) observer is presented. It is based on a relationship between the geometries of a surface in a scene and its occluding contour: If the viewing direction of the observer is along a principal direction for a surface point whose projection is on the contour, surface shape (i.e., curvature) at the surface point can be recovered from the contour. An observer that purposefully changes viewpoint in order to achieve a well-defined geometric relationship with respect to a 3D shape prior to its recognition is used. It is shown that there is a simple and efficient viewing strategy that allows the observer to align the viewing direction with one of the two principal directions for a point on the surface. Experimental results demonstrate that the method can be easily implemented and can provide reliable shape information.> Kiriakos N. Kutulakos, Charles R. Dyer |
CVPR | 1 |
| 1992 | Fast Computation of the Euclidian Distance Maps for Binary ImagesabstractA simple algorithm is given for the computation of the Euclidian distance from the set of black points in an N × N black and white image, for all points in the image. The running time is O(N2 log N) and O(N) extra space is required. The algorithm is suitable for implementation on a parallel machine. Mihail N. Kolountzakis, Kiriakos N. Kutulakos |
Inf. Process. Lett. | 2 |