VLDB 2026 Research / reviewers in the wild / expert
Ramesh Raskar
dblp:r/RameshRaskar
· DBLP profile ↗
185ranked-venue papers
17as first author
34since 2021 · last 2025
0000-0002-3254-3224ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 140 · 15 first-author · 21 since 2021Artificial intelligence and machine learning · 75 · 2 first-author · 24 since 2021Human-computer interaction and ubiquitous computing · 23 · 6 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Computer networks · 4 · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Co-Dream: Collaborative Dream Synthesis over Decentralized ModelsabstractFederated Learning (FL) has pioneered the idea of "share wisdom not raw data" to enable collaborative learning over decentralized data. FL achieves this goal by averaging model parameters instead of centralizing data. However, representing "wisdom" in the form of model parameters has its own limitations including the requirement for uniform model architectures across clients and communication overhead proportional to model size. In this work we introduce Co-Dream a framework for representing "wisdom" in data space instead of model parameters. Here, clients collaboratively optimize random inputs based on their locally trained models and aggregate gradients of their inputs. Our proposed approach overcomes the aforementioned limitations and comes with additional benefits such as adaptive optimization and interpretable representation of knowledge. We empirically demonstrate the effectiveness of Co-Dream and compare its performance with existing techniques. Abhishek Singh 0005, Gauri Gupta, Yichuan Shi, Alex Dang, Ritvik Kapila, Sheshank Shankar, Mohammed Ehab, Ramesh Raskar |
AAAI | 8 |
| 2025 | Blurred LiDAR for Sharper 3D: Robust Handheld 3D Scanning with Diffuse LiDAR and RGBabstract3D surface reconstruction is essential across applications of virtual reality, robotics, and mobile scanning. However, RGB-based reconstruction often fails in low-texture, low-light, and low-albedo scenes. Handheld LiDARs, now common on mobile devices, aim to address these challenges by capturing depth information from time-of-flight measurements of a coarse grid of projected dots. Yet, these sparse LiDARs struggle with scene coverage on limited input views, leaving large gaps in depth information. In this work, we propose using an alternative class of "blurred" LiDAR that emits a diffuse flash, greatly improving scene coverage but introducing spatial ambiguity from mixed time-of-flight measurements across a wide field of view. To handle these ambiguities, we propose leveraging the complementary strengths of diffuse LiDAR with RGB. We introduce a Gaussian surfel-based rendering framework with a scene-adaptive loss function that dynamically balances RGB and diffuse LiDAR signals. We demonstrate that, surprisingly, diffuse LiDAR can outperform traditional sparse LiDAR, enabling robust 3D scanning with accurate color and geometry estimation in challenging environments. Nikhil Behari, Aaron Young, Siddharth Somasundaram, Tzofi Klinghoffer, Akshat Dave, Ramesh Raskar |
CVPR | 6 |
| 2025 | Enhancing Autonomous Navigation by Imaging Hidden Objects Using Single-Photon LiDARabstractRobust autonomous navigation in environments with limited visibility remains a critical challenge in robotics. We present a novel approach that leverages Non-Line-of-Sight (NLOS) sensing using single-photon LiDAR to improve visibility and enhance autonomous navigation. Our method enables mobile robots to “see around corners” by utilizing multi-bounce light information, effectively expanding their perceptual range without additional infrastructure. We propose a three-module pipeline: (1) Sensing, which captures multi-bounce histograms using SPAD-based LiDAR; (2) Perception, which estimates occupancy maps of hidden regions from these histograms using a convolutional neural network; and (3) Control, which allows a robot to follow safe paths based on the estimated occupancy. We evaluate our approach through simulations and real-world experiments on a mobile robot navigating an L-shaped corridor with hidden obstacles. Our work represents the first experimental demonstration of NLOS imaging for autonomous navigation, paving the way for safer and more efficient robotic systems operating in complex environments. We also contribute a novel dynamics-integrated transient rendering framework for simulating NLOS scenarios, facilitating future research in this domain. Aaron Young, Nevindu Batagoda, Harry Zhang, Akshat Dave, Adithya Kumar Pediredla, Dan Negrut, Ramesh Raskar |
ICRA | 7 |
| 2025 | On the Limits of Agency in Agent-based Models
Ayush Chopra, Nurullah Giray Kuru, Ramesh Raskar, Arnau Quera-Bofarull |
AAMAS | 4 |
| 2025 | Shoot-Bounce-3D: Single-Shot Occlusion-Aware 3D from Lidar by Decomposing Two-Bounce Lightabstract3D scene reconstruction from a single measurement is challenging, especially in the presence of occluded regions and specular materials, such as mirrors. We address these challenges by leveraging single-photon lidars. These lidars estimate depth from light that is emitted into the scene and reflected directly back to the sensor. However, they can also measure light that bounces multiple times in the scene before reaching the sensor. This multi-bounce light contains additional information that can be used to recover dense depth, occluded geometry, and material properties. Prior work with single-photon lidar, however, has only demonstrated these use cases when a laser sequentially illuminates one scene point at a time. We instead focus on the more practical – and challenging – scenario of illuminating multiple scene points simultaneously. The complexity of light transport due to the combined effects of multiplexed illumination, two-bounce light, shadows, and specular reflections is challenging to invert analytically. Instead, we propose a data-driven method to invert light transport in single-photon lidar. To enable this approach, we create the first large-scale simulated dataset of ~100k lidar transients for indoor scenes. We use this dataset to learn a prior on complex light transport, enabling measured two-bounce light to be decomposed into the constituent contributions from each laser spot. Finally, we experimentally demonstrate how this decomposed light can be used to infer 3D geometry in scenes with occlusions and mirrors from a single measurement. Our code and dataset are released on our project webpage. Tzofi Klinghoffer, Siddharth Somasundaram, Xiaoyu Xiang, Yuchen Fan 0001, Christian Richardt, Akshat Dave, Ramesh Raskar |
SIGGRAPH Asia | 7 |
| 2025 | Event Cameras Meet SPADs for High-Speed, Low-Bandwidth ImagingabstractTraditional cameras face a trade-off between low-light performance and high-speed imaging: longer exposure times to capture sufficient light results in motion blur, whereas shorter exposures result in Poisson-corrupted noisy images. While burst photography techniques help mitigate this tradeoff, conventional cameras are fundamentally limited in their sensor noise characteristics. Event cameras and single-photon avalanche diode (SPAD) sensors have emerged as promising alternatives to conventional cameras due to their desirable properties. SPADs are capable of single-photon sensitivity with microsecond temporal resolution, and event cameras can measure brightness changes up to 1 MHz with low bandwidth requirements. We show that these properties are complementary, and can help achieve low-light, high-speed image reconstruction with low bandwidth requirements. We introduce a sensor fusion framework to combine SPADs with event cameras to improve the reconstruction of high-speed, low-light scenes while reducing the high bandwidth cost associated with using every SPAD frame. Our evaluation, on both synthetic and real sensor data, demonstrates significant enhancements ($> 5$>5 dB PSNR) in reconstructing low-light scenes at high temporal resolution (100 kHz) compared to conventional cameras. Event-SPAD fusion shows great promise for real-world applications, such as robotics or medical imaging. Manasi Muglikar, Siddharth Somasundaram, Akshat Dave, Edoardo Charbon, Ramesh Raskar, Davide Scaramuzza 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | NeST: Neural Stress Tensor Tomography by leveraging 3D PhotoelasticityabstractPhotoelasticity enables full-field stress analysis in transparent objects through stress-induced birefringence. Existing techniques are limited to two-dimensional (2D) slices and require destructively slicing the object. Recovering the internal three-dimensional (3D) stress distribution of the entire object is challenging, as it involves solving a tensor tomography problem and handling phase wrapping ambiguities. We introduce NeST, an analysis-by-synthesis approach for reconstructing 3D stress tensor fields as neural implicit representations from polarization measurements. Our key insight is to jointly handle phase unwrapping and tensor tomography using a differentiable forward model based on Jones calculus. Our non-linear model faithfully matches real captures, unlike prior linear approximations. We develop an experimental multi-axis polariscope setup to capture 3D photoelasticity and experimentally demonstrate that NeST reconstructs the internal stress distribution for objects with varying shape and force conditions. Additionally, we showcase novel applications in stress analysis, such as visualizing photoelastic fringes by virtually slicing the object and viewing photoelastic fringes from unseen viewpoints. NeST paves the way for scalable non-destructive 3D photoelastic analysis. Akshat Dave, Aaron Young, Ramesh Raskar, Wolfgang Heidrich, Ashok Veeraraghavan |
ACM Trans. Graph. | 4 |
| 2024 | PlatoNeRF: 3D Reconstruction in Plato's Cave via Single-View Two-Bounce Lidarabstract3D reconstruction from a single-view is challenging because of the ambiguity from monocular cues and lack of information about occluded regions. Neural radiance fields (NeRF), while popular for view synthesis and 3D reconstruction, are typically reliant on multi-view images. Existing methods for single-view 3D reconstruction with NeRF rely on either data priors to hallucinate views of occluded regions, which may not be physically accurate, or shadows observed by RGB cameras, which are difficult to detect in ambient light and low albedo backgrounds. We propose using time-of-flight data captured by a single-photon avalanche diode to overcome these limitations. Our method models two-bounce optical paths with NeRF, using lidar transient data for supervision. By leveraging the advantages of both NeRF and two-bounce light measured by lidar, we demonstrate that we can reconstruct visible and occluded geometry without data priors or reliance on controlled ambient lighting or scene albedo. In addition, we demonstrate improved generalization under practical constraints on sensor spatial- and temporal-resolution. We believe our method is a promising direction as single-photon lidars become ubiquitous on consumer devices, such as phones, tablets, and headsets. Tzofi Klinghoffer, Xiaoyu Xiang, Siddharth Somasundaram, Yuchen Fan 0001, Christian Richardt, Ramesh Raskar |
CVPR | 6 |
| 2024 | SIMBA: Split Inference - Mechanisms, Benchmarks and Attacks
Abhishek Singh 0005, Vivek Sharma 0001, Rohan Sukumaran, John Mose, Jeffrey Chiu, Justin Yu, Ramesh Raskar |
ECCV (76) | 7 |
| 2024 | DecentNeRFs: Decentralized Neural Radiance Fields from Crowdsourced Images
Zaid Tasneem, Akshat Dave, Abhishek Singh 0005, Kushagra Tiwary, Praneeth Vepakomma, Ashok Veeraraghavan, Ramesh Raskar |
ECCV (59) | 7 |
| 2024 | Handheld Mapping of Specular Surfaces Using Consumer-Grade Flash LiDARabstractWe propose an approach to leverage multi-bounce returns of a flash LiDAR on portable smartphones for 3D specular surface reconstruction. Traditional LiDAR systems assume that all returns are one-bounce returns, which can lead to an overestimation of the true mirror surface and cause it to appear as if there is a hole. However, in reality, returns from mirror surfaces follow multi-bounce paths. We operate with a consumer-grade, coarse multi-beam flash LiDAR, enabling real-time mapping on an affordable and portable smartphone. To address the challenges posed by the coarse setup, where the transmitter and receiver are co-located, we propose solving the association problem using the ‘reciprocal pair’ algorithm. This algorithm can distinguish between different types of bounces from multi-bounce returns. We have demonstrated detection over multiple consecutive frames for dense mirror mapping. In addition to 3D reconstruction, we show that multi-bounce returns enhance performance in applications such as segmentation and novel view synthesis. Our method can be integrated with state-of-the-art learned-based models, enhancing their robustness in discerning ambiguous scenarios. Importantly, our approach can map various specular surfaces like mirrors and glasses without assuming specific shapes, and it can operate on non-perpendicular specular-diffuse surface pairs. Tsung-Han Lin, Connor Henley, Siddharth Somasundaram, Akshat Dave, Moshe Laifenfeld, Ramesh Raskar |
ICCP | 6 |
| 2024 | Incentive-Aware Federated Learning with Training-Time Model RewardsabstractIn federated learning (FL), incentivizing contributions of training resources (e.g., data, compute) from potentially competitive clients is crucial. Existing incentive mechanisms often distribute post-training monetary rewards, which suffer from practical challenges of timeliness and feasibility of the rewards. Rewarding the clients after the completion of training may incentivize them to abort the collaboration, and monetizing the contribution is challenging in practice. To address these problems, we propose an incentive-aware algorithm that offers differentiated training-time model rewards for each client at each FL iteration. We theoretically prove that such a $\textit{local}$ design ensures the $\textit{global}$ objective of client incentivization. Through theoretical analyses, we further identify the issue of error propagation in model rewards and thus propose a stochastic reference-model recovery strategy to ensure theoretically that all the clients eventually obtain the optimal model in the limit. We perform extensive experiments to demonstrate the superior incentivizing performance of our method compared to existing baselines. Zhaoxuan Wu, Mohammad Mohammadi Amiri, Ramesh Raskar, Kian Hsiang Low |
ICLR | 3 |
| 2024 | Data Acquisition via Experimental Design for Data MarketsabstractThe acquisition of training data is crucial for machine learning applications. Data markets can increase the supply of data, particularly in data-scarce domains such as healthcare, by incentivizing potential data providers to join the market. A major challenge for a data buyer in such a market is choosing the most valuable data points from a data seller. Unlike prior work in data valuation, which assumes centralized data access, we propose a federated approach to the data acquisition problem that is inspired by linear experimental design. Our proposed data acquisition method achieves lower prediction error without requiring labeled validation data and can be optimized in a fast and federated procedure. The key insight of our work is that a method that directly estimates the benefit of acquiring data for test set prediction is particularly compatible with a decentralized market setting. Charles Lu 0001, Baihe Huang, Sai Praneeth Karimireddy, Praneeth Vepakomma, Michael I. Jordan, Ramesh Raskar |
NeurIPS | 6 |
| 2024 | Mix2SFL: Two-Way Mixup for Scalable, Accurate, and Communication-Efficient Split Federated LearningabstractIn recent years, split learning (SL) has emerged as a promising distributed learning framework that can utilize big data in parallel without privacy leakage while reducing client-side computing resources. In the initial implementation of SL, however, the server serves multiple clients sequentially incurring high latency. Parallel implementation of SL can alleviate this latency problem, but existing Parallel SL algorithms compromise scalability due to its fundamental structural problem. To this end, our previous works have proposed two scalable Parallel SL algorithms, dubbed SGLR and LocFedMix-SL, by solving the aforementioned fundamental problem of the Parallel SL structure. In this article, we propose a novel Parallel SL framework, coined Mix2SFL, that can ameliorate both accuracy and communication-efficiency while still ensuring scalability. Mix2SFL first supplies more samples to the server through a manifold mixup between the smashed data uploaded to the server as in SmashMix of LocFedMix-SL, and then averages the split-layer gradient as in GradMix of SGLR, followed by local model aggregation as in SFL. Numerical evaluation corroborates that Mix2SFL achieves improved performance in both accuracy and latency compared to the state-of-the-art SL algorithm with scalability guarantees. Moreover, its convergence speed as well as privacy guarantee are validated through the experimental results. Seungeun Oh, Hyelin Nam, Jihong Park, Praneeth Vepakomma, Ramesh Raskar, Mehdi Bennis, Seong-Lyun Kim |
IEEE Trans. Big Data | 5 |
| 2023 | Fundamentals of Task-Agnostic Data ValuationabstractWe study valuing the data of a data owner/seller for a data seeker/buyer. Data valuation is often carried out for a specific task assuming a particular utility metric, such as test accuracy on a validation set, that may not exist in practice. In this work, we focus on task-agnostic data valuation without any validation requirements. The data buyer has access to a limited amount of data (which could be publicly available) and seeks more data samples from a data seller. We formulate the problem as estimating the differences in the statistical properties of the data at the seller with respect to the baseline data available at the buyer. We capture these statistical differences through second moment by measuring diversity and relevance of the seller’s data for the buyer; we estimate these measures through queries to the seller without requesting the raw data. We design the queries with the proposed approach so that the seller is blind to the buyer’s raw data and has no knowledge to fabricate responses to the queries to obtain a desired outcome of the diversity and relevance trade-off. We will show through extensive experiments on real tabular and image datasets that the proposed estimates capture the diversity and relevance of the seller’s data for the buyer. Mohammad Mohammadi Amiri, Frédéric Berdoz, Ramesh Raskar |
AAAI | 3 |
| 2023 | Role of Transients in Two-Bounce Non-Line-of-Sight ImagingabstractThe goal of non-line-of-sight (NLOS) imaging is to image objects occluded from the camera's field of view using multiply scattered light. Recent works have demonstrated the feasibility of two-bounce (2B) NLOS imaging by scanning a laser and measuring cast shadows of occluded objects in scenes with two relay surfaces. In this work, we study the role of time-of-flight (ToF) measurements, i.e. transients, in 2B-NLOS under multiplexed illumination. Specifically, we study how ToF information can reduce the number of measurements and spatial resolution needed for shape reconstruction. We present our findings with respect to tradeoffs in (1) temporal resolution, (2) spatial resolution, and (3) number of image captures by studying SNR and recoverability as functions of system parameters. This leads to a formal definition of the mathematical constraints for 2B lidar. We believe that our work lays an analytical ground- workfor design of future NLOS imaging systems, especially as ToF sensors become increasingly ubiquitous. Siddharth Somasundaram, Akshat Dave, Connor Henley, Ashok Veeraraghavan, Ramesh Raskar |
CVPR | 5 |
| 2023 | ORCa: Glossy Objects as Radiance-Field CamerasabstractReflections on glossy objects contain valuable and hidden information about the surrounding environment. By converting these objects into cameras, we can unlock exciting applications, including imaging beyond the camera's field-of-view and from seemingly impossible vantage points, e.g. from reflections on the human eye. However, this task is challenging because reflections depend jointly on object geometry, material properties, the 3D environment, and the observer's viewing direction. Our approach converts glossy objects with unknown geometry into radiance-field cameras to image the world from the object's perspective. Our key insight is to convert the object surface into a virtual sensor that captures cast reflections as a 2D projection of the 5D environment radiance field visible to and surrounding the object. We show that recovering the environment radiance fields enables depth and radiance estimation from the object to its surroundings in addition to beyond field-of-view novel-view synthesis, i.e. rendering of novel views that are only directly visible to the glossy object present in the scene, but not the observer. Moreover, using the radiance field we can image around occluders caused by close-by objects in the scene. Our method is trained end-to-end on multi-view images of the object and jointly estimates object geometry, diffuse radiance, and the 5D environment radiance field. For more information, visit our website. Kushagra Tiwary, Akshat Dave, Nikhil Behari, Tzofi Klinghoffer, Ashok Veeraraghavan, Ramesh Raskar |
CVPR | 6 |
| 2023 | Towards Viewpoint Robustness in Bird's Eye View SegmentationabstractAutonomous vehicles (AV) require that neural networks used for perception be robust to different viewpoints if they are to be deployed across many types of vehicles without the repeated cost of data collection and labeling for each. AV companies typically focus on collecting data from diverse scenarios and locations, but not camera rig configurations, due to cost. As a result, only a small number of rig variations exist across most fleets. In this paper, we study how AV perception models are affected by changes in camera viewpoint and propose a way to scale them across vehicle types without repeated data collection and labeling. Using bird’s eye view (BEV) segmentation as a motivating task, we find through extensive experiments that existing perception models are surprisingly sensitive to changes in camera viewpoint. When trained with data from one camera rig, small changes to pitch, yaw, depth, or height of the camera at inference time lead to large drops in performance. We introduce a technique for novel view synthesis and use it to transform collected data to the viewpoint of target rigs, allowing us to train BEV segmentation models for diverse target rigs without any additional data collection or labeling cost. To analyze the impact of viewpoint changes, we leverage synthetic data to mitigate other gaps (content, ISP, etc). Our approach is then trained on real data and evaluated on synthetic data, enabling evaluation on diverse target rigs. We release all data for use in future work. Our method is able to recover an average of 14.7% of the IoU that is otherwise lost when deploying to new rigs. Tzofi Klinghoffer, Jonah Philion, Wenzheng Chen, Or Litany, Zan Gojcic, Jungseock Joo, Ramesh Raskar, Sanja Fidler, José M. Álvarez 0004 |
ICCV | 7 |
| 2023 | DISeR: Designing Imaging Systems with Reinforcement LearningabstractImaging systems consist of cameras to encode visual information about the world and perception models to interpret this encoding. Cameras contain (1) illumination sources, (2) optical elements, and (3) sensors, while perception models use (4) algorithms. Directly searching over all combinations of these four building blocks to design an imaging system is challenging due to the size of the search space. Moreover, cameras and perception models are often designed independently, leading to sub-optimal task performance. In this paper, we formulate these four building blocks of imaging systems as a context-free grammar (CFG), which can be automatically searched over with a learned camera designer to jointly optimize the imaging system with task-specific perception models. By transforming the CFG to a state-action space, we then show how the camera designer can be implemented with reinforcement learning to intelligently search over the combinatorial space of possible imaging system configurations. We demonstrate our approach on two tasks, depth estimation and camera rig design for autonomous vehicles, showing that our method yields rigs that outperform industry-wide standards. We believe that our proposed approach is an important step towards automating imaging system design. Our project page is https://tzofi.github.io/diser. Tzofi Klinghoffer, Kushagra Tiwary, Nikhil Behari, Bhavya Agrawalla, Ramesh Raskar |
ICCV | 5 |
| 2023 | Federated Conformal Predictors for Distributed Uncertainty QuantificationabstractConformal prediction is emerging as a popular paradigm for providing rigorous uncertainty quantification in machine learning since it can be easily applied as a post-processing step to already trained models. In this paper, we extend conformal prediction to the federated learning setting. The main challenge we face is data heterogeneity across the clients --- this violates the fundamental tenet of *exchangeability* required for conformal prediction. We propose a weaker notion of *partial exchangeability*, better suited to the FL setting, and use it to develop the Federated Conformal Prediction (FCP) framework. We show FCP enjoys rigorous theoretical guarantees and excellent empirical performance on several computer vision and medical imaging datasets. Our results demonstrate a practical approach to incorporating meaningful uncertainty quantification in distributed and heterogeneous environments. We provide code used in our experiments https://github.com/clu5/federated-conformal. Charles Lu 0001, Yaodong Yu, Sai Praneeth Karimireddy, Michael I. Jordan, Ramesh Raskar |
ICML | 5 |
| 2023 | Posthoc privacy guarantees for collaborative inference with modified Propose-Test-ReleaseabstractCloud-based machine learning inference is an emerging paradigm where users query by sending their data through a service provider who runs an ML model on that data and returns back the answer. Due to increased concerns over data privacy, recent works have proposed Collaborative Inference (CI) to learn a privacy-preserving encoding of sensitive user data before it is shared with an untrusted service provider. Existing works so far evaluate the privacy of these encodings through empirical reconstruction attacks. In this work, we develop a new framework that provides formal privacy guarantees for an arbitrarily trained neural network by linking its local Lipschitz constant with its local sensitivity. To guarantee privacy using local sensitivity, we extend the Propose-Test-Release (PTR) framework to make it tractable for neural network queries. We verify the efficacy of our framework experimentally on real-world datasets and elucidate the role of Adversarial Representation Learning (ARL) in improving the privacy-utility trade-off. Abhishek Singh 0005, Praneeth Vepakomma, Vivek Sharma 0001, Ramesh Raskar |
NeurIPS | 4 |
| 2023 | SAFE-PASS: Stewardship, Advocacy, Fairness and Empowerment in Privacy, Accountability, Security, and Safety for Vulnerable GroupsabstractOur vision is to achieve societally responsible secure and trustworthy cyberspace that puts algorithmic and technological checks and balances on the indiscriminate sharing and analysis of data. We achieve this vision in a holistic manner by framing research directions with four major considerations: (i) Expanding knowledge and understanding of security and privacy perceptions and expectations in vulnerable groups, which significantly contribute to their unwillingness to share data, and use that knowledge to drive research in (a) mitigating missing/imbalanced data problems, (b) understanding and modeling security and privacy risks of data sharing, and (c) modeling utility of data sharing. (ii) Developing a risk-adaptive, policy model capable of capturing and articulating security and privacy expectations of users that are relevant in a particular context and develops associated technology to ensure provenance and accountability. (iii) Developing robust AI/ML algorithms that are transparent and explainable with respect to fairness and bias to reduce/eliminate discrimination, misuse, privacy violations, or other cyber-crimes. (iv) Developing models and techniques for a nuanced, contextually adaptive, and graded privacy paradigm that allows trade-offs between privacy and utility. Towards this, in this paper we present the SAFE-PASS framework to provide Stewardship, Advocacy, Fairness and Empowerment in Privacy, Accountability, Security, and Safety for Vulnerable Groups. Indrajit Ray, Bhavani Thuraisingham, Jaideep Vaidya, Sharad Mehrotra, Vijayalakshmi Atluri, Indrakshi Ray, Murat Kantarcioglu, Ramesh Raskar, Babak Salimi, Steven J. Simske, Nalini Venkatasubramanian, Vivek K. Singh 0001 |
SACMAT | 8 |
| 2022 | PrivateMail: Supervised Manifold Learning of Deep Features with Privacy for Image RetrievalabstractDifferential Privacy offers strong guarantees such as immutable privacy under any post-processing. In this work, we propose a differentially private mechanism called PrivateMail for performing supervised manifold learning. We then apply it to the use case of private image retrieval to obtain nearest matches to a client’s target image from a server’s database. PrivateMail releases the target image as part of a differentially private manifold embedding. We give bounds on the global sensitivity of the manifold learning map in order to obfuscate and release embeddings with differential privacy inducing noise. We show that PrivateMail obtains a substantially better performance in terms of the privacy-utility trade off in comparison to several baselines on various datasets. We share code for applying PrivateMail at http://tiny.cc/PrivateMail. Praneeth Vepakomma, Julia Balla, Ramesh Raskar |
AAAI | 3 |
| 2022 | Blind Inference: An Automated Privacy-Preserving Prediction Service using Secure Multi-Party Computation for Medical Applications
Gharib Gharibi, Babak Poorebrahim Gilkalaye, Praneeth Vepakomma, Zachi Attia, Riddhiman Das, Suraj Kapa, Ramesh Raskar |
AMIA | 7 |
| 2022 | Learning to Censor by Noisy Sampling
Ayush Chopra, Abhinav Java, Abhishek Singh 0005, Vivek Sharma 0001, Ramesh Raskar |
ECCV (13) | 5 |
| 2022 | Decouple-and-Sample: Protecting Sensitive Information in Task Agnostic Data Release
Abhishek Singh 0005, Ethan Garza, Ayush Chopra, Praneeth Vepakomma, Vivek Sharma 0001, Ramesh Raskar |
ECCV (13) | 6 |
| 2022 | Towards Learning Neural Representations from Shadows
Kushagra Tiwary, Tzofi Klinghoffer, Ramesh Raskar |
ECCV (33) | 3 |
| 2022 | Physics vs. Learned Priors: Rethinking Camera and Algorithm Design for Task-Specific ImagingabstractCameras were originally designed using physics-based heuristics to capture aesthetic images. In recent years, there has been a transformation in camera design from being purely physics-driven to increasingly data-driven and task-specific. In this paper, we present a framework to understand the building blocks of this nascent field of end-to-end design of camera hardware and algorithms. As part of this framework, we show how methods that exploit both physics and data have become prevalent in imaging and computer vision, underscoring a key trend that will continue to dominate the future of task-specific camera design. Finally, we share current barriers to progress in end-to-end design, and hypothesize how these barriers can be overcome. Tzofi Klinghoffer, Siddharth Somasundaram, Kushagra Tiwary, Ramesh Raskar |
ICCP | 4 |
| 2022 | An Automated Framework for Distributed Deep Learning-A Tool DemoabstractSplit learning (SL) is a distributed deep-learning approach that enables individual data owners to train a shared model over their joint data without exchanging it with one another. SL has been the subject of much research in recent years, leading to the development of several versions for facilitating distributed learning. However, the majority of this work mainly focuses on optimizing the training process while largely ignoring the design and implementation of practical tool support. To fill this gap, we present our automated software framework for training deep neural networks from decentralized data based on our extended version of SL, termed Blind Learning. Specifically, we shed light on the underlying optimization algorithm, explain the design and implementation details of our framework, and present our preliminary evaluation results. We demonstrate that Blind Learning is 65% more computationally efficient than SL and can produce better performing models. Moreover, we show that running the same job in our framework is at least 4.5× faster than PySyft. Our goal is to spur the development of proper tool support for distributed deep learning. Gharib Gharibi, Anissa Khan, Babak Poorebrahim Gilkalaye, Praneeth Vepakomma, Ramesh Raskar, Steve Penrod, Greg Storm, Riddhiman Das |
ICDCS | 6 |
| 2022 | LocFedMix-SL: Localize, Federate, and Mix for Improved Scalability, Convergence, and Latency in Split LearningabstractSplit learning (SL) is a promising distributed learning framework that enables to utilize the huge data and parallel computing resources of mobile devices. SL is built upon a model-split architecture, wherein a server stores an upper model segment that is shared by different mobile clients storing its lower model segments. Without exchanging raw data, SL achieves high accuracy and fast convergence by only uploading smashed data from clients and downloading global gradients from the server. Nonetheless, the original implementation of SL sequentially serves multiple clients, incurring high latency with many clients. A parallel implementation of SL has great potential in reducing latency, yet existing parallel SL algorithms resort to compromising scalability and/or convergence speed. Motivated by this, the goal of this article is to develop a scalable parallel SL algorithm with fast convergence and low latency. As a first step, we identify that the fundamental bottleneck of existing parallel SL comes from the model-split and parallel computing architectures, under which the server-client model updates are often imbalanced, and the client models are prone to detach from the server’s model. To fix this problem, by carefully integrating local parallelism, federated learning, and mixup augmentation techniques, we propose a novel parallel SL framework, coined LocFedMix-SL. Simulation results corroborate that LocFedMix-SL achieves improved scalability, convergence speed, and latency, compared to sequential SL as well as the state-of-the-art parallel SL algorithms such as SplitFed and LocSplitFed. Seungeun Oh, Jihong Park, Praneeth Vepakomma, Sihun Baek, Ramesh Raskar, Mehdi Bennis, Seong-Lyun Kim |
WWW | 5 |
| 2021 | DISCO: Dynamic and Invariant Sensitive Channel Obfuscation for Deep Neural NetworksabstractRecent deep learning models have shown remarkable performance in image classification. While these deep learning systems are getting closer to practical deployment, the common assumption made about data is that it does not carry any sensitive information. This assumption may not hold for many practical cases, especially in the domain where an individual’s personal information is involved, like healthcare and facial recognition systems. We posit that selectively removing features in this latent space can protect the sensitive information and provide better privacy-utility trade-off. Consequently, we propose DISCO which learns a dynamic and data driven pruning filter to selectively obfuscate sensitive information in the feature space. We propose diverse attack schemes for sensitive inputs & attributes and demonstrate the effectiveness of DISCO against state-of-the-art methods through quantitative and qualitative evaluation. Finally, we also release an evaluation benchmark dataset of 1 million sensitive representations to encourage rigorous exploration of novel attack and defense schemes at https://github.com/splitlearning/InferenceBenchmark. Abhishek Singh 0005, Ayush Chopra, Ethan Garza, Emily Zhang, Praneeth Vepakomma, Vivek Sharma 0001, Ramesh Raskar |
CVPR | 7 |
| 2021 | NoPeek-Infer: Preventing face reconstruction attacks in distributed inference after on-premise trainingabstractFor models trained on-premise but deployed in a distributed fashion across multiple entities, we demonstrate that minimizing distance correlation between sensitive data such as faces and intermediary representations enables prediction while preventing reconstruction attacks. Leakage (measured using distance correlation between input and intermediate representations) is the risk associated with the reconstruction of raw face data from intermediary representations that are communicated in a distributed setting. We demonstrate on face datasets that our method is resilient to reconstruction attacks during distributed inference while maintaining information required to sustain good classification accuracy. We share modular code for performing NoPeek-Infer at http://tiny.cc/nopeek along with corresponding trained models for benchmarking attack techniques. Praneeth Vepakomma, Abhishek Singh 0005, Emily Zhang, Otkrist Gupta, Ramesh Raskar |
FG | 5 |
| 2021 | AirMixML: Over-the-Air Data Mixup for Inherently Privacy-Preserving Edge Machine LearningabstractWireless channels can be inherently privacy preserving by distorting the received signals due to channel noise, and superpositioning multiple signals over-the-air. By harnessing these natural distortions and superpositions by wireless channels, we propose a novel privacy-preserving machine learning (ML) framework at the network edge, coined over-the-air mixup ML (AirMixML). In AirMixML, multiple workers transmit analog-modulated signals of their private data samples to an edge server who trains an ML model using the received noisy-and-superpositioned samples. AirMixML coincides with model training using mixup data augmentation achieving comparable accuracy to that with raw data samples. From a privacy perspective, AirMixML is a differentially private (DP) mechanism limiting the disclosure of each worker's private sample information at the server, while the worker's transmit power determines the privacy disclosure level. To this end, we develop a fractional channel-inversion power control (PC) method, a-Dirichlet mixup PC (DirMix(a)-PC), wherein for a given global power scaling factor after channel inversion, each worker's local power contribution to the superpositioned signal is controlled by the Dirichlet dispersion ratio a. Mathematically, we derive a closed-form expression clarifying the relationship between the local and global PC factors to guarantee a target DP level. By simulations, we provide DirMix(α)-PC design guidelines to improve accuracy, privacy, and energy-efficiency. Finally, AirMixML with DirMix(a)-PC is shown to achieve reasonable accuracy compared to a privacy-violating baseline with neither superposition nor PC. Yusuke Koda, Jihong Park, Mehdi Bennis, Praneeth Vepakomma, Ramesh Raskar |
GLOBECOM | 5 |
| 2021 | Objects as Cameras: Estimating High-Frequency Illumination from ShadowsabstractWe recover high-frequency information encoded in the shadows cast by an object to estimate a hemispherical photograph from the viewpoint of the object, effectively turning objects into cameras. Estimating environment maps is useful for advanced image editing tasks such as relighting, object insertion or removal, and material parameter estimation. Because the problem is ill-posed, recent works in illumination recovery have tackled the problem of low- frequency lighting for object insertion, rely upon specular surface materials, or make use of data-driven methods that are susceptible to hallucination without physically plausible constraints. We incorporate an optimization scheme to update scene parameters that could enable practical capture of real-world scenes. Furthermore, we develop a methodology for evaluating expected recovery performance for different types and shapes of objects. Tristan Swedish, Connor Henley, Ramesh Raskar |
ICCV | 3 |
| 2020 | Deep Polarization Cues for Transparent Object SegmentationabstractSegmentation of transparent objects is a hard, open problem in computer vision. Transparent objects lack texture of their own, adopting instead the texture of scene background. This paper reframes the problem of transparent object segmentation into the realm of light polarization, i.e., the rotation of light waves. We use a polarization camera to capture multi-modal imagery and couple this with a unique deep learning backbone for processing polarization input data. Our method achieves instance segmentation on cluttered, transparent objects in various scene and background conditions, demonstrating an improvement over traditional image-based approaches. As an application we use this for robotic bin picking of transparent objects. Agastya Kalra, Vage Taamazyan, Supreeth Krishna Rao, Kartik Venkataraman, Ramesh Raskar, Achuta Kadambi |
CVPR | 5 |
| 2020 | Imaging Behind Occluders Using Two-Bounce Light
Connor Henley, Tomohiro Maeda, Tristan Swedish, Ramesh Raskar |
ECCV (29) | 4 |
| 2020 | 3d Imaging For Thermal Cameras Using Structured LightabstractOptical 3D sensing technologies are exploited for many applications in autonomous vehicles, manufacturing, and consumer products. However, existing techniques may suffer in certain challenging conditions, where scattering may occur due to particles. While the light in the visible and near IR spectrum is affected by scattering, long-wave IR (LWIR) tends to experience less scattering, especially when the particles are much smaller than the incident radiation. We propose and demonstrate the expansion of structured light scanning approaches into the LWIR spectrum using a thermal camera and black body radiation source. We then validate the results produced against ground truth scans from traditional structured light scanners. Additional means for projecting these scanning patterns are also discussed alongside potential drawbacks and challenges of this technique associated with future adoption. Jack Erdozain, Kazuto Ichimaru, Tomohiro Maeda, Hiroshi Kawasaki, Ramesh Raskar, Achuta Kadambi |
ICIP | 5 |
| 2019 | Thermal Non-Line-of-Sight ImagingabstractWe propose a novel non-line-of-sight (NLOS) imaging framework with long-wave infrared (IR). At long-wave IR wavelengths, certain physical parameters are more favorable for high-fidelity reconstruction. In contrast to prior work in visible light NLOS, at long-wave IR wavelengths, the hidden heat source acts as a light source. This simplifies the problem to a single bounce problem. In addition, surface reflectance has a much stronger specular reflection in the long-wave IR spectrum than in the visible light spectrum. We reformulate a light transport model that leverages these favorable physical properties of long-wave IR. Specifically, we demonstrate 2D shape recovery and 3D localization of a hidden object. Furthermore, we demonstrate near real-time and robust NLOS pose estimation of a human figure, the first such demonstration, to our knowledge. Tomohiro Maeda, Ramesh Raskar, Achuta Kadambi |
ICCP | 3 |
| 2019 | Multi-Velocity Neural Networks for Facial Expression Recognition in VideosabstractWe present a new action recognition deep neural network which adaptively learns the best action velocities in addition to the classification. While deep neural networks have reached maturity for image understanding tasks, we are still exploring network topologies and features to handle the richer environment of video clips. Here, we tackle the problem of multiple velocities in action recognition, and provide state-of-the-art results for facial expression recognition, on known and new collected datasets. We further provide the training steps for our semi-supervised network, suited to learn from huge unlabeled datasets with only a fraction of labeled examples. Otkrist Gupta, Dan Raviv, Ramesh Raskar |
IEEE Trans. Affect. Comput. | 3 |
| 2018 | Pairwise Confusion for Fine-Grained Visual Classification
Abhimanyu Dubey, Otkrist Gupta, Pei Guo, Ramesh Raskar, Ryan Farrell, Nikhil Naik 0003 |
ECCV (12) | 4 |
| 2018 | Unlimited Sampling of Sparse SignalsabstractIn a recent paper [1], we introduced the concept of “Unlimited Sampling”. This unique approach circumvents the clipping or saturation problem in conventional analog-to-digital converters (ADCs) by considering a radically different ADC architecture which resets the input voltage before saturation. Such ADCs, also known as Self-Reset ADCs (SR-ADCs), allow for sensing modulo samples. In analogy to Shannon's sampling theorem, the unlimited sampling theorem proves that a bandlimited signal can be recovered from modulo samples provided that a certain sampling density criterion, that is independent of the ADC threshold, is satisfied. In this way, our result allows for perfect recovery of a bandlimited function whose amplitude exceeds the ADC threshold by orders of magnitude. By capitalizing on this result, in this paper, we consider the inverse problem of recovering a sparse signal from its low-pass filtered version. This problem frequently arises in several areas of science and engineering and in context of signal processing, it is studied in several flavors, namely, sparse or FRI sampling, super-resolution and sparse deconvolution. By considering the SR-ADC architecture, we develop a sampling theory for modulo sampling of lowpass filtered spikes. Our main result consists of a new sparse sampling theorem and an algorithm which stably recovers a K -sparse signal from low-pass, modulo samples. We validate our results using numerical experiments. Ayush Bhandari, Felix Krahmer, Ramesh Raskar |
ICASSP | 3 |
| 2018 | Dynamic heterodyne interferometryabstractDynamic interferometry enables snapshot recovery of phase images by using polarization phase shifting. However, the phase estimate is susceptible to influence from sources of ambient light having uncontrolled polarization. We present a novel method, dynamic heterodyne interferometry (DHI), as a means to mitigate phase bias from ambient light sources, while retaining dynamic potential. Tomohiro Maeda, Achuta Kadambi, Yoav Y. Schechner, Ramesh Raskar |
ICCP | 4 |
| 2018 | Towards photography through realistic fogabstractImaging through fog has important applications in industries such as self-driving cars, augmented driving, airplanes, helicopters, drones and trains. Here we show that time profiles of light reflected from fog have a distribution (Gamma) that is different from light reflected from objects occluded by fog (Gaussian). This helps to distinguish between background photons reflected from the fog and signal photons reflected from the occluded object. Based on this observation, we recover reflectance and depth of a scene obstructed by dense, dynamic, and heterogeneous fog. For practical use cases, the imaging system is designed in optical reflection mode with minimal footprint and is based on LIDAR hardware. Specifically, we use a single photon avalanche diode (SPAD) camera that time-tags individual detected photons. A probabilistic computational framework is developed to estimate the fog properties from the measurement itself, without prior knowledge. Other solutions are based on radar that suffers from poor resolution (due to the long wavelength), or on time gating that suffers from low signal-to-noise ratio. The suggested technique is experimentally evaluated in a wide range of fog densities created in a fog chamber It demonstrates recovering objects 57cm away from the camera when the visibility is 37cm. In that case it recovers depth with a resolution of 5cm and scene reflectance with an improvement of 4dB in PSNR and 3.4× reconstruction quality in SSIM over time gating techniques. Guy Satat, Matthew Tancik, Ramesh Raskar |
ICCP | 3 |
| 2018 | Unlimited Sampling of Sparse Sinusoidal MixturesabstractIn parallel to Shannon's sampling theorem, the recent theory of unlimited sampling yields that a bandlimited function with high dynamic range can be recovered exactly from oversampled, low dynamic range samples. In this way, the unlimited sampling methodology circumvents the dynamic range problem that limits the use of conventional analog-to-digital converters (ADCs) which are prone to clipping or saturation problem. The unlimited sampling theorem is made practicable by using a unique ADC architecture-the self-reset ADC or the SR-ADC-which resets voltage before clipping, thus producing modulo or wrapped samples. While retaining full dynamic range of the input signal, surprisingly, the sampling density prescribed by the unlimited sampling theorem is independent of the maximum recordable voltage of the new ADC and depends only on the signal bandwidth. As the corresponding problem of signal recovery from such modulo samples arises in various applications with different signal models, where the original result does not directly apply, the original paper continues to trigger research follow-ups. In this paper, we investigate the case of sampling and reconstruction of a mixture of K sinusoids from such modulo samples. This problem is at the heart of spectral estimation theory and application areas include active sensing, ranging, source localization, interferometry and direction-of-arrival estimation. By relying on the SR-ADCs, we develop a method for recovery of K-sparse, sum-of-sinusoids from finitely many wrapped samples, thus avoiding clipping or saturation. As our signal model is completely characterized by K pairs of amplitudes and frequencies, we obtain a parametric sampling theorem; we complement it with a recovery algorithm. Numerical demonstrations validate the effectivity of our approach. Ayush Bhandari, Felix Krahmer, Ramesh Raskar |
ISIT | 3 |
| 2018 | Maximum-Entropy Fine Grained ClassificationabstractFine-Grained Visual Classification (FGVC) is an important computer vision problem that involves small diversity within the different classes, and often requires expert annotators to collect data. Utilizing this notion of small visual diversity, we revisit Maximum-Entropy learning in the context of fine-grained classification, and provide a training routine that maximizes the entropy of the output probability distribution for training convolutional neural networks on FGVC tasks. We provide a theoretical as well as empirical justification of our approach, and achieve state-of-the-art performance across a variety of classification tasks in FGVC, that can potentially be extended to any fine-tuning task. Our method is robust to different hyperparameter values, amount of training data and amount of training label noise and can hence be a valuable tool in many similar problems. Abhimanyu Dubey, Otkrist Gupta, Ramesh Raskar, Nikhil Naik 0003 |
NeurIPS | 3 |
| 2018 | Distributed learning of deep neural network over multiple agents
Otkrist Gupta, Ramesh Raskar |
J. Netw. Comput. Appl. | 2 |
| 2018 | Illumination invariants in deep video expression recognition
Otkrist Gupta, Dan Raviv, Ramesh Raskar |
Pattern Recognit. | 3 |
| 2017 | Zensei: Embedded, Multi-electrode Bioimpedance Sensing for Implicit, Ubiquitous User RecognitionabstractInteractions and connectivity is increasingly expanding to shared objects and environments, such as furniture, vehicles, lighting, and entertainment systems. For transparent personalization in such contexts, we see an opportunity for embedded recognition, to complement traditional, explicit authentication. We introduce Zensei, an implicit sensing system that leverages bio-sensing, signal processing and machine learning to classify uninstrumented users by their body's electrical properties. Zensei could allow many objects to recognize users. E.g., phones that unlock when held, cars that automatically adjust mirrors and seats, or power tools that restore user settings. We introduce wide-spectrum bioimpedance hardware that measures both amplitude and phase. It extends previous approaches through multi-electrode sensing and high-speed wireless data collection for embedded devices. We implement the sensing in devices and furniture, where unique electrode configurations generate characteristic profiles based on user's unique electrical properties. Finally, we discuss results from a comprehensive longitudinal 22-day data collection experiment with 46 subjects. Our analysis shows promising classification accuracy and low false acceptance rate. Munehiko Sato, Rohan S. Puri, Alex Olwal, Yosuke Ushigome, Lukas Franciszkiewicz, Deepak Chandra, Ivan Poupyrev, Ramesh Raskar |
CHI | 8 |
| 2017 | Sampling without time: Recovering echoes of light via temporal phase retrievalabstractThis paper considers the problem of sampling and reconstruction of a continuous-time sparse signal without assuming the knowledge of the sampling instants or the sampling rate. This topic has its roots in the problem of recovering multiple echoes of light from its low-pass filtered and auto-correlated, time-domain measurements. Our work is closely related to the topic of sparse phase retrieval and in this context, we discuss the advantage of phase-free measurements. While this problem is ill-posed, cues based on physical constraints allow for its appropriate regularization. We validate our theory with experiments based on customized, optical time-of-flight imaging sensors. What singles out our approach is that our sensing method allows for temporal phase retrieval as opposed to the usual case of spatial phase retrieval. Preliminary experiments and results demonstrate a compelling capability of our phaseretrieval based imaging device. Ayush Bhandari, Aurélien Bourquard, Ramesh Raskar |
ICASSP | 3 |
| 2017 | Learning Gaze Transitions from Depth to Improve Video Saliency EstimationabstractIn this paper we introduce a novel Depth-Aware Video Saliency approach to predict human focus of attention when viewing videos that contain a depth map (RGBD) on a 2D screen. Saliency estimation in this scenario is highly important since in the near future 3D video content will be easily acquired yet hard to display. Despite considerable progress in 3D display technologies, most are still expensive and require special glasses for viewing, so RGBD content is primarily viewed on 2D screens, removing the depth channel from the final viewing experience. We train a generative convolutional neural network that predicts the 2D viewing saliency map for a given frame using the RGBD pixel values and previous fixation estimates in the video. To evaluate the performance of our approach, we present a new comprehensive database of 2D viewing eye-fixation ground-truth for RGBD videos. Our experiments indicate that it is beneficial to integrate depth into video saliency estimates for content that is viewed on a 2D display. We demonstrate that our approach outperforms state-of-the-art methods for video saliency, achieving 15% relative improvement. George Leifman, Dmitry Rudoy, Tristan Swedish, Eduardo Bayro-Corrochano, Ramesh Raskar |
ICCV | 5 |
| 2017 | Designing Neural Network Architectures using Reinforcement Learning
Bowen Baker, Otkrist Gupta, Nikhil Naik 0003, Ramesh Raskar |
ICLR (Poster) | 4 |
| 2017 | Depth Sensing Using Geometrically Constrained Polarization Normals
Achuta Kadambi, Vage Taamazyan, Boxin Shi, Ramesh Raskar |
Int. J. Comput. Vis. | 4 |
| 2017 | LRA: Local Rigid Averaging of Stretchable Non-rigid Shapes
Dan Raviv, Eduardo Bayro-Corrochano, Ramesh Raskar |
Int. J. Comput. Vis. | 3 |
| 2016 | Macroscopic Interferometry: Rethinking Depth Estimation with Frequency-Domain Time-of-FlightabstractA form of meter-scale, macroscopic interferometry is proposed using conventional time-of-flight (ToF) sensors. Today, ToF sensors use phase-based sampling, where the phase delay between emitted and received, high-frequency signals encodes distance. This paper examines an alternative ToF architecture, inspired by micron-scale, microscopic interferometry, that relies only on frequency sampling: we refer to our proposed macroscopic technique as Frequency-Domain Time of Flight (FD-ToF). The proposed architecture offers several benefits over existing phase ToF systems, such as robustness to phase wrapping and implicit resolution of multi-path interference, all while capturing the same number of subframes. A prototype camera is constructed to demonstrate macroscopic interferometry at meter scale. Achuta Kadambi, Jamie Schiel, Ramesh Raskar |
CVPR | 3 |
| 2016 | Deep Learning the City: Quantifying Urban Perception at a Global Scale
Abhimanyu Dubey, Nikhil Naik 0003, Devi Parikh, Ramesh Raskar, César A. Hidalgo 0001 |
ECCV (1) | 4 |
| 2016 | Time-resolved image demixingabstractWhen multiple light paths combine at a given location on an image sensor, an image mixture is created. Demixing or recovering the original constituent components in such cases is a highly ill-posed problem. A number of elegant solutions have thus been developed in the literature, relying on measurement diversity such as polarization, shift, motion, or scene features. In this paper, we approach the image-mixing problem as a time-resolved phenomenon-if every photon arriving at the sensor could be time-stamped, the demixing problem would then amount to separating transient events in time. Based on this idea, we first show that, while acquiring measurements is prohibitive and challenging in the time domain, this task is surprisingly straightforward in the frequency domain. We then establish a link between frequency-domain measurements and consumer time-of-flight (ToF) imaging. Finally, we propose a demixing algorithm, relying only on magnitude information of the ToF sensor. We show that our problem is closely tied to the topic of phase retrieval and that for K-image mixture, (K2-K)/2+1 magnitude-only ToF measurements suffice to demix images exactly in noiseless settings. Our developments are corroborated with experiments on synthetic and ToF data acquired using the Microsoft Kinect sensor. Ayush Bhandari, Aurélien Bourquard, Shahram Izadi, Ramesh Raskar |
ICASSP | 4 |
| 2016 | Super-resolved time-of-flight sensing via FRI sampling theoryabstractOptical time-of-flight (ToF) sensors can measure scene depth accurately by projection and reception of an optical signal. The range to a surface in the path of the emitted signal is proportional to the delay time of the light echo or the reflected signal. In practice, a diverging beam may be subject to multi-echo backscatter, and all these echoes must be resolved to estimate the multiple depths. In this paper, we propose a method for super-resolution of optical ToF signals. Our contributions are twofold. Starting with a general image formation model common to most ToF sensors, we draw a striking analogy of ToF systems with sampling theory. Based on our model, we reformulate the ToF super-resolution problem as a parameter estimation problem pivoted around the finite-rate-of-innovation framework. In particular, we show that super-resolution of multi-echo backscattered signal amounts to recovery of Dirac impulses from low-pass measurements. Our theory is corroborated by analysis of data collected from a photon counting, LiDAR sensor, showing the effectiveness of our non-iterative and computationally efficient algorithm. Ayush Bhandari, Andrew M. Wallace, Ramesh Raskar |
ICASSP | 3 |
| 2016 | Tensor low-rank and sparse light field photography
Mahdad Hosseini Kamal, Barmak Heshmat, Ramesh Raskar, Pierre Vandergheynst, Gordon Wetzstein |
Comput. Vis. Image Underst. | 3 |
| 2016 | Occluded Imaging with Time-of-Flight SensorsabstractWe explore the question of whether phase-based time-of-flight (TOF) range cameras can be used for looking around corners and through scattering diffusers. By connecting TOF measurements with theory from array signal processing, we conclude that performance depends on two primary factors: camera modulation frequency and the width of the specular lobe (“shininess”) of the wall. For purely Lambertian walls, commodity TOF sensors achieve resolution on the order of meters between targets. For seemingly diffuse walls, such as posterboard, the resolution is drastically reduced, to the order of 10cm. In particular, we find that the relationship between reflectance and resolution is nonlinear—a slight amount of shininess can lead to a dramatic improvement in resolution. Since many realistic scenes exhibit a slight amount of shininess, we believe that off-the-shelf TOF cameras can look around corners. Achuta Kadambi, Hang Zhao 0021, Boxin Shi, Ramesh Raskar |
ACM Trans. Graph. | 4 |
| 2015 | SpecTrans: Versatile Material Classification for Interaction with Textureless, Specular and Transparent SurfacesabstractSurface and object recognition is of significant importance in ubiquitous and wearable computing. While various techniques exist to infer context from material properties and appearance, they are typically neither designed for real-time applications nor for optically complex surfaces that may be specular, textureless, and even transparent. These materials are, however, becoming increasingly relevant in HCI for transparent displays, interactive surfaces, and ubiquitous computing. We present SpecTrans, a new sensing technology for surface classification of exotic materials, such as glass, transparent plastic, and metal. The proposed technique extracts optical features by employing laser and multi-directional, multi-spectral LED illumination that leverages the material's optical properties. The sensor hardware is small in size, and the proposed classification method requires significantly lower computational cost than conventional image-based methods, which use texture features or reflectance analysis, thereby providing real-time performance for ubiquitous computing. Our evaluation of the sensing technique for nine different transparent materials, including air, shows a promising recognition rate of 99.0%. We demonstrate a variety of possible applications using SpecTrans' capabilities. Munehiko Sato, Shigeo Yoshida, Alex Olwal, Boxin Shi, Atsushi Hiyama, Tomohiro Tanikawa, Michitaka Hirose, Ramesh Raskar |
CHI | 8 |
| 2015 | A light transport model for mitigating multipath interference in Time-of-flight sensorsabstractContinuous-wave Time-of-flight (TOF) range imaging has become a commercially viable technology with many applications in computer vision and graphics. However, the depth images obtained from TOF cameras contain scene dependent errors due to multipath interference (MPI). Specifically, MPI occurs when multiple optical reflections return to a single spatial location on the imaging sensor. Many prior approaches to rectifying MPI rely on sparsity in optical reflections, which is an extreme simplification. In this paper, we correct MPI by combining the standard measurements from a TOF camera with information from direct and global light transport. We report results on both simulated experiments and physical experiments (using the Kinect sensor). Our results, evaluated against ground truth, demonstrate a quantitative improvement in depth accuracy. Nikhil Naik 0003, Achuta Kadambi, Christoph Rhemann, Shahram Izadi, Ramesh Raskar, Sing Bing Kang |
CVPR | 5 |
| 2015 | Super-resolution in Phase SpaceabstractThis work considers the problem of super-resolution. The goal is to resolve a Dirac distribution from knowledge of its discrete, low-pass, Fourier measurements. Classically, such problems have been dealt with parameter estimation methods. Recently, it has been shown that convex-optimization based formulations facilitate a continuous time solution to the super-resolution problem. Here we treat super-resolution from low-pass measurements in Phase Space. The Phase Space transformation parametrically generalizes a number of well known unitary mappings such as the Fractional Fourier, Fresnel, Laplace and Fourier transforms. Consequently, our work provides a general super-resolution strategy which is backward compatible with the usual Fourier domain result. We consider low-pass measurements of Dirac distributions in Phase Space and show that the super-resolution problem can be cast as Total Variation minimization. Remarkably, even though are setting is quite general, the bounds on the minimum separation distance of Dirac distributions is comparable to existing methods. Ayush Bhandari, Yonina C. Eldar, Ramesh Raskar |
ICASSP | 3 |
| 2015 | Unbounded High Dynamic Range Photography Using a Modulo CameraabstractThis paper presents a novel framework to extend the dynamic range of images called Unbounded High Dynamic Range (UHDR) photography with a modulo camera. A modulo camera could theoretically take unbounded radiance levels by keeping only the least significant bits. We show that with limited bit depth, very high radiance levels can be recovered from a single modulus image with our newly proposed unwrapping algorithm for natural images. We can also obtain an HDR image with details equally well preserved for all radiance levels by merging the least number of modulus images. Synthetic experiment and experiment with a real modulo camera show the effectiveness of the proposed approach. Hang Zhao 0021, Boxin Shi, Christy Fernandez-Cull, Sai-Kit Yeung, Ramesh Raskar |
ICCP | 5 |
| 2015 | Polarized 3D: High-Quality Depth Sensing with Polarization CuesabstractCoarse depth maps can be enhanced by using the shape information from polarization cues. We propose a framework to combine surface normals from polarization (hereafter polarization normals) with an aligned depth map. Polarization normals have not been used for depth enhancement before. This is because polarization normals suffer from physics-based artifacts, such as azimuthal ambiguity, refractive distortion and fronto-parallel signal degradation. We propose a framework to overcome these key challenges, allowing the benefits of polarization to be used to enhance depth maps. Our results demonstrate improvement with respect to state-of-the-art 3D reconstruction techniques. Achuta Kadambi, Vage Taamazyan, Boxin Shi, Ramesh Raskar |
ICCV | 4 |
| 2015 | Depth Map Estimation and Colorization of Anaglyph Images Using Local Color Prior and Reverse Intensity DistributionabstractIn this paper, we present a joint iterative anaglyph stereo matching and colorization framework for obtaining a set of disparity maps and colorized images. Conventional stereo matching algorithms fail when addressing anaglyph images that do not have similar intensities on their two respective view images. To resolve this problem, we propose two novel data costs using local color prior and reverse intensity distribution factor for obtaining accurate depth maps. To colorize an anaglyph image, each pixel in one view is warped to another view using the obtained disparity values of non-occluded regions. A colorization algorithm using optimization is then employed with additional constraint to colorize the remaining occluded regions. Experimental results confirm that the proposed unified framework is robust and produces accurate depth maps and colorized stereo images. Williem 0001, Ramesh Raskar, In Kyu Park |
ICCV | 2 |
| 2015 | Extreme Computational PhotographyabstractThe Camera Culture Group at the MIT Media Lab aims to create a new class of imaging platforms. This talk will discuss three tracks of research: femto photography, retinal imaging, and 3D displays. Femto Photography consists of femtosecond laser illumination, picosecond-accurate detectors and mathematical reconstruction techniques allowing researchers to visualize propagation of light. Direct recording of reflected or scattered light at such a frame rate with sufficient brightness is nearly impossible. Using an indirect 'stroboscopic' method that records millions of repeated measurements by careful scanning in time and viewpoints we can rearrange the data to create a 'movie' of a nanosecond long event. Femto photography and a new generation of nano-photography (using ToF cameras) allow powerful inference with computer vision in presence of scattering. EyeNetra is a mobile phone attachment that allows users to test their own eyesight. The device reveals corrective measures thus bringing vision to billions of people who would not have had access otherwise. Another project, eyeMITRA, is a mobile retinal imaging solution that brings retinal exams to the realm of routine care, by lowering the cost of the imaging device to a 10th of its current cost and integrating the device with image analysis software and predictive analytics. This provides early detection of Diabetic Retinopathy that can change the arc of growth of the world's largest cause of blindness. Finally the talk will describe novel lightfield cameras and lightfield displays that require a compressive optical architecture to deal with high bandwidth requirements of 4D signals Ramesh Raskar |
UIST | 1 |
| 2015 | Relativistic Effects for Time-Resolved Light TransportabstractAbstract We present a real‐time framework which allows interactive visualization of relativistic effects for time‐resolved light transport. We leverage data from two different sources: real‐world data acquired with an effective exposure time of less than 2 picoseconds, using an ultra‐fast imaging technique termed femto‐photography, and a transient renderer based on ray‐tracing. We explore the effects of time dilation, light aberration, frequency shift and radiance accumulation by modifying existing models of these relativistic effects to take into account the time‐resolved nature of light propagation. Unlike previous works, we do not impose limiting constraints in the visualization, allowing the virtual camera to explore freely a reconstructed 3D scene depicting dynamic illumination. Moreover, we consider not only linear motion, but also acceleration and rotation of the camera. We further introduce, for the first time, a pinhole camera model into our relativistic rendering framework, and account for subsequent changes in focal length and field of view as the camera moves through the scene. Adrián Jarabo, Belén Masiá, Andreas Velten, Christopher Barsi, Ramesh Raskar, Diego Gutierrez |
Comput. Graph. Forum | 5 |
| 2015 | Scale Invariant Metrics of Volumetric DatasetsabstractNature reveals itself in similar structures of different scales. A child and an adult share similar organs yet dramatically differ in size. Comparing the two is a challenging task to a computerized approach as scale and shape are coupled. Recently, it was shown that a local measure based on the Gaussian curvature can be used to normalize the local metric of a surface and then to extract global features and distances. In this paper we consider higher dimensions; specifically, we construct a scale invariant metric for volumetric domains which can be used in analysis of medical datasets such as computed tomography (CT) and magnetic resonance imaging (MRI). Dan Raviv, Ramesh Raskar |
SIAM J. Imaging Sci. | 2 |
| 2015 | eyeSelfie: self directed eye alignment using reciprocal eye box imagingabstractEye alignment to the optical system is very critical in many modern devices, such as for biometrics, gaze tracking, head mounted displays, and health. We show alignment in the context of the most difficult challenge: retinal imaging. Alignment in retinal imaging, even conducted by a physician, is very challenging due to precise alignment requirements and lack of direct user eye gaze control. Self-imaging of the retina is nearly impossible. We frame this problem as a user-interface (UI) challenge. We can create a better UI by controlling the eye box of a projected cue. Our key concept is to exploit the reciprocity, "If you see me, I see you", to develop near eye alignment displays. Two technical aspects are critical: a) tightness of the eye box and (b) the eye box discovery comfort. We demonstrate that previous pupil forming display architectures are not adequate to address alignment in depth. We then analyze two ray-based designs to determine efficacious fixation patterns. These ray based displays and a sequence of user steps allow lateral (x, y) and depth (z) wise alignment to deal with image centering and focus. We show a highly portable prototype and demonstrate the effectiveness through a user study. Tristan Swedish, Karin Roesch, Ikhyun Lee, Krishna Rastogi, Shoshana Bernstein, Ramesh Raskar |
ACM Trans. Graph. | 6 |
| 2014 | Sub-pixel Layout for Super-Resolution with Images in the Octic Group
Boxin Shi, Hang Zhao 0021, Moshe Ben-Ezra, Sai-Kit Yeung, Christy Fernandez-Cull, R. Hamilton Shepard, Christopher Barsi, Ramesh Raskar |
ECCV (1) | 8 |
| 2014 | Sparse Linear Operator identification without sparse regularization? Applications to mixed pixel problem in Time-of-Flight/Range imagingabstractIn this paper, we consider the problem of Sparse Linear Operator identification which is also linked with the topic of Sparse Deconvolution. In its abstract form, the problem can be stated as follows: Given a well behaved probing function, is it possible to identify a Sparse Linear Operator from its response to the function? We present a constructive solution to this problem. Furthermore, our approach is devoid of any sparsity inducing penalty term and explores the idea of parametric modeling. Consequently, our algorithm is non-iterative by design and circumvents tuning of any regularization parameter. Our approach is computationally efficient when compared the ℓ0/ℓ1-norm regularized counterparts. Our work addresses a problem of industrial significance: decomposition of mixed-pixels in Time-of-Flight/Range imaging. In this case, each pixel records range measurements from multiple contributing depths and the goal is to isolate each depth. Practical experiments corroborate our theoretical set-up and establish the efficiency of our approach, that is, speed-up in processing with lesser mean squared error. We also derive Cramér-Rao Bounds for performance characterization. Ayush Bhandari, Achuta Kadambi, Ramesh Raskar |
ICASSP | 3 |
| 2014 | A switchable light field camera architecture with Angle Sensitive Pixels and dictionary-based sparse codingabstractWe propose a flexible light field camera architecture that is at the convergence of optics, sensor electronics, and applied mathematics. Through the co-design of a sensor that comprises tailored, Angle Sensitive Pixels and advanced reconstruction algorithms, we show that—contrary to light field cameras today—our system can use the same measurements captured in a single sensor image to recover either a high-resolution 2D image, a low-resolution 4D light field using fast, linear processing, or a high-resolution light field using sparsity-constrained optimization. Matthew Hirsch, Sriram Sivaramakrishnan, Suren Jayasuriya, Albert Wang 0004, Alyosha C. Molnar, Ramesh Raskar, Gordon Wetzstein |
ICCP | 6 |
| 2014 | Demultiplexing illumination via low cost sensing and nanosecond codingabstractSeveral computer vision algorithms require a sequence of photographs taken in different illumination conditions, which has spurred development in the area of illumination multiplexing. Various techniques for optimizing the multiplexing process already exist, but are geared toward regular or high speed cameras. Such cameras are fast, but code on the order of milliseconds. In this paper we propose a fusion of two popular contexts, time of flight range cameras and illumination multiplexing. Time of flight cameras are a low cost, consumer-oriented technology capable of acquiring range maps at 30 frames per second. Such cameras have a natural connection to conventional illumination multiplexing strategies as both paradigms rely on the capture of multiple shots and synchronized illumination. While previous work on illumination multiplexing has exploited coding at millisecond intervals, we repurpose sensors that are ordinarily used in time of flight imaging to demultiplex via nanosecond coding strategies. Achuta Kadambi, Ayush Bhandari, Refael Whyte, Adrian A. Dorrington, Ramesh Raskar |
ICCP | 5 |
| 2014 | Skin perfusion photographyabstractThe separation of global and direct light components of a scene is highly useful for scene analysis, as each component offers different information about illumination-scene-detector interactions. Relying on ray optics, the technique is important in computational photography, but it is often under appreciated in the biomedical imaging community, where wave interference effects are utilized. Nevertheless, such coherent optical systems lend themselves naturally to global-direct separation methods because of the high spatial frequency nature of speckle interference patterns. Here, we extend global-direct separation to laser speckle contrast imaging (LSCI) system to reconstruct speed maps of blood flow in skin. We compare experimental results with a speckle formation model of moving objects and show that the reconstructed map of skin perfusion is improved over the conventional case. Guy Satat, Christopher Barsi, Ramesh Raskar |
ICCP | 3 |
| 2014 | Phase messaging method for time-of-flight camerasabstractUbiquitous light emitting devices and low-cost commercial digital cameras facilitate optical wireless communication system such as visual MIMO where handheld cameras communicate with electronic displays. While intensity-based optical communications are more prevalent in camera-display messaging, we present a novel method that uses modulated light phase for messaging and time-of-flight (ToF) cameras for receivers. With intensity-based methods, light signals can be degraded by reflections and ambient illumination. By comparison, communication using ToF cameras is more robust against challenging lighting conditions. Additionally, the concept of phase messaging can be combined with intensity messaging for a significant data rate advantage. In this work, we design and construct a phase messaging array (PMA), which is the first of its kind, to communicate to a ToF depth camera by manipulating the phase of the depth camera's infrared light signal. The array enables message variation spatially using a plane of infrared light emitting diodes and temporally by varying the induced phase shift. In this manner, the phase messaging array acts as the transmitter by electronically controlling the light signal phase. The ToF camera acts as the receiver by observing and recording a time-varying depth. We show a complete implementation of a 3×3 prototype array with custom hardware and demonstrating average bit accuracy as high as 97.8%. The prototype data rate with this approach is 1 Kbps that can be extended to approximately 10 Mbps. Wenjia Yuan, Richard E. Howard, Kristin J. Dana, Ramesh Raskar, Ashwin Ashok, Marco Gruteser, Narayan B. Mandayam |
ICCP | 4 |
| 2014 | Computational Schlieren Photography with Light Field Probes
Gordon Wetzstein, Wolfgang Heidrich, Ramesh Raskar |
Int. J. Comput. Vis. | 3 |
| 2014 | Decomposing Global Light Transport Using Time of Flight Imaging
Di Wu 0006, Andreas Velten, Matthew O'Toole, Belén Masiá, Amit K. Agrawal, Qionghai Dai, Ramesh Raskar |
Int. J. Comput. Vis. | 7 |
| 2014 | Ultra-fast Lensless Computational Imaging through 5D Frequency Analysis of Time-resolved Light Transport
Di Wu 0006, Gordon Wetzstein, Christopher Barsi, Thomas Willwacher, Qionghai Dai, Ramesh Raskar |
Int. J. Comput. Vis. | 6 |
| 2014 | A compressive light field projection systemabstractFor about a century, researchers and experimentalists have strived to bring glasses-free 3D experiences to the big screen. Much progress has been made and light field projection systems are now commercially available. Unfortunately, available display systems usually employ dozens of devices making such setups costly, energy inefficient, and bulky. We present a compressive approach to light field synthesis with projection devices. For this purpose, we propose a novel, passive screen design that is inspired by angle-expanding Keplerian telescopes. Combined with high-speed light field projection and nonnegative light field factorization, we demonstrate that compressive light field projection is possible with a single device. We build a prototype light field projector and angle-expanding screen from scratch, evaluate the system in simulation, present a variety of results, and demonstrate that the projector can alternatively achieve super-resolved and high dynamic range 2D image display when used with a conventional screen. Matthew Hirsch, Gordon Wetzstein, Ramesh Raskar |
ACM Trans. Graph. | 3 |
| 2014 | Eyeglasses-free display: towards correcting visual aberrations with computational light field displaysabstractMillions of people worldwide need glasses or contact lenses to see or read properly. We introduce a computational display technology that predistorts the presented content for an observer, so that the target image is perceived without the need for eyewear. By designing optics in concert with prefiltering algorithms, the proposed display architecture achieves significantly higher resolution and contrast than prior approaches to vision-correcting image display. We demonstrate that inexpensive light field displays driven by efficient implementations of 4D prefiltering algorithms can produce the desired vision-corrected imagery, even for higher-order aberrations that are difficult to be corrected with glasses. The proposed computational display architecture is evaluated in simulation and with a low-cost prototype device. Fu-Chung Huang, Gordon Wetzstein, Brian A. Barsky, Ramesh Raskar |
ACM Trans. Graph. | 4 |
| 2014 | Toward BxDF display using multilayer diffractionabstractWith a wide range of applications in product design and optical watermarking, computational BxDF display has become an emerging trend in the graphics community. In this paper, we analyze the design space of BxDF displays and show that existing approaches cannot reproduce arbitrary BxDFs. In particular, existing surface-based fabrication techniques are often limited to generating only specific angular frequencies, angle-shift-invariant radiance distributions, and sometimes only symmetric BxDFs. To overcome these limitations, we propose diffractive multilayer BxDF displays. We derive forward and inverse methods to synthesize patterns that are printed on stacked, high-resolution transparencies and reproduce prescribed BxDFs with unprecedented degrees of freedom within the limits of available fabrication techniques. Ramesh Raskar |
ACM Trans. Graph. | 1 |
| 2013 | CrowdCam: Instantaneous Navigation of Crowd Images Using Angled GraphabstractWe present a near real-time algorithm for interactively exploring a collectively captured moment without explicit 3D reconstruction. Our system favors immediacy and local coherency to global consistency. It is common to represent photos as vertices of a weighted graph, where edge weights measure similarity or distance between pairs of photos. We introduce Angled Graphs as a new data structure to organize collections of photos in a way that enables the construction of visually smooth paths. Weighted angled graphs extend weighted graphs with angles and angle weights which penalize turning along paths. As a result, locally straight paths can be computed by specifying a photo and a direction. The weighted angled graphs of photos used in this paper can be regarded as the result of discretizing the Riemannian geometry of the high dimensional manifold of all possible photos. Ultimately, our system enables everyday people to take advantage of each others' perspectives in order to create on-the-spot spatiotemporal visual experiences similar to the popular bullet-time sequence. We believe that this type of application will greatly enhance shared human experiences spanning from events as personal as parents watching their children's football game to highly publicized red carpet galas. Aydin Arpa, Luca Ballan, Rahul Sukthankar, Gabriel Taubin, Marc Pollefeys, Ramesh Raskar |
3DV | 6 |
| 2013 | 8D: interacting with a relightable glasses-free 3D displayabstractWe present an 8-dimensional (8D) display that allows glasses-free viewing of 3D imagery, whist capturing and reacting to incident environmental and user controlled light sources. We demonstrate two interactive possibilities enabled by our lens-array-based hardware prototype, and realtime GPU-accelerated software pipeline. Additionally, we describe a path to deploying such displays in the future, using current Sensor-in-Pixel (SIP) LCD panels, which physically collocate sensing and display elements. Matthew Hirsch, Shahram Izadi, Henry Holtzman, Ramesh Raskar |
CHI | 4 |
| 2013 | Discovering the Structure of a Planar Mirror System from Multiple Observations of a Single PointabstractWe investigate the problem of identifying the position of a viewer inside a room of planar mirrors with unknown geometry in conjunction with the room's shape parameters. We consider the observations to consist of angularly resolved depth measurements of a single scene point that is being observed via many multi-bounce interactions with the specular room geometry. Applications of this problem statement include areas such as calibration, acoustic echo cancelation and time-of-flight imaging. We theoretically analyze the problem and derive sufficient conditions for a combination of convex room geometry, observer, and scene point to be reconstruct able. The resulting constructive algorithm is exponential in nature and, therefore, not directly applicable to practical scenarios. To counter the situation, we propose theoretically devised geometric constraints that enable an efficient pruning of the solution space and develop a heuristic randomized search algorithm that uses these constraints to obtain an effective solution. We demonstrate the effectiveness of our algorithm on extensive simulations as well as in a challenging real-world calibration scenario. Ilya Reshetouski, Alkhazur Manakov, Ayush Bhandari, Ramesh Raskar, Hans-Peter Seidel, Ivo Ihrke |
CVPR | 4 |
| 2013 | Simultaneous geometry and texture display based on lateral force for touchscreenabstractDisplaying haptic and tactile information on a touchscreen is one of the key technologies for human interfaces. In this paper we propose a method that allows the user to simultaneously feel both large bump and small textures through a screen. Our method employs lateral force and direction-controlled mechanical vibration (around 0-400 Hz). The technology allows not only geometrical shapes but also textures to be felt. Our experiments revealed that with our proposed method almost all people can simultaneously detect geometry and texture information on a touchscreen. In addition, providing adequate direction for the vibration enhances the feelings of softness. Satoshi Saga, Ramesh Raskar |
World Haptics | 2 |
| 2013 | Coded focal stack photographyabstractWe present coded focal stack photography as a computational photography paradigm that combines a focal sweep and a coded sensor readout with novel computational algorithms. We demonstrate various applications of coded focal stacks, including photography with programmable non-planar focal surfaces and multiplexed focal stack acquisition. By leveraging sparse coding techniques, coded focal stacks can also be used to recover a full-resolution depth and all-in-focus (AIF) image from a single photograph. Coded focal stack photography is a significant step towards a computational camera architecture that facilitates high-resolution post-capture refocusing, flexible depth of field, and 3D imaging. Jin-Li Suo, Gordon Wetzstein, Qionghai Dai, Ramesh Raskar |
ICCP | 5 |
| 2013 | High-rank coded aperture projection for extended depth of fieldabstractProjectors require large apertures to maximize light throughput. Unfortunately, this leads to shallow depths of field (DOF), hence blurry images, when projecting on non-planar surfaces, such as cultural heritage sites, curved screens, or when sharing visual information in everyday environments. We introduce high-rank coded aperture projectors - a new computational display technology that combines optical designs with computational processing to overcome depth of field limitations of conventional devices. In particular, we employ high-speed spatial light modulators (SLMs) on the image plane and in the aperture of modified projectors. The patterns displayed on these SLMs are computed with a new mathematical framework that uses high-rank light field factorizations and directly exploits the limited temporal resolution and contrast sensitivity of the human visual system. With an experimental prototype projector, we demonstrate significantly increased DOF as compared to conventional technology. Chenguang Ma, Jin-Li Suo, Qionghai Dai, Ramesh Raskar, Gordon Wetzstein |
ICCP | 4 |
| 2013 | Display adaptive 3D content remapping
Belén Masiá, Gordon Wetzstein, Carlos Aliaga, Ramesh Raskar, Diego Gutierrez |
Comput. Graph. | 4 |
| 2013 | Near-invariant blur for depth and 2D motion via time-varying light field analysisabstractRecently, several camera designs have been proposed for either making defocus blur invariant to scene depth or making motion blur invariant to object motion. The benefit of such invariant capture is that no depth or motion estimation is required to remove the resultant spatially uniform blur. So far, the techniques have been studied separately for defocus and motion blur, and object motion has been assumed 1D (e.g., horizontal). This article explores a more general capture method that makes both defocus blur and motion blur nearly invariant to scene depth and in-plane 2D object motion. We formulate the problem as capturing a time-varying light field through a time-varying light field modulator at the lens aperture, and perform 5D (4D light field + 1D time) analysis of all the existing computational cameras for defocus/motion-only deblurring and their hybrids. This leads to a surprising conclusion that focus sweep, previously known as a depth-invariant capture method that moves the plane of focus through a range of scene depth during exposure, is near-optimal both in terms of depth and 2D motion invariance and in terms of high-frequency preservation for certain combinations of depth and motion ranges. Using our prototype camera, we demonstrate joint defocus and motion deblurring for moving scenes with depth variation. Yosuke Bando, Henry Holtzman, Ramesh Raskar |
ACM Trans. Graph. | 3 |
| 2013 | Adaptive image synthesis for compressive displaysabstractRecent years have seen proposals for exciting new computational display technologies that arecompressivein the sense that they generate high resolution images or light fields with relatively few display parameters. Image synthesis for these types of displays involves two major tasks: sampling and rendering high-dimensional target imagery, such as light fields or time-varying light fields, as well as optimizing the display parameters to provide a good approximation of the target content. In this paper, we introduce an adaptive optimization framework for compressive displays that generates high quality images and light fields using only a fraction of the total plenoptic samples. We demonstrate the framework for a large set of display technologies, including several types of auto-stereoscopic displays, high dynamic range displays, and high-resolution displays. We achieve significant performance gains, and in some cases are able to process data that would be infeasible with existing methods. Felix Heide, Gordon Wetzstein, Ramesh Raskar, Wolfgang Heidrich |
ACM Trans. Graph. | 3 |
| 2013 | Coded time of flight cameras: sparse deconvolution to address multipath interference and recover time profilesabstractTime of flight cameras produce real-time range maps at a relatively low cost using continuous wave amplitude modulation and demodulation. However, they are geared to measure range (or phase) for a single reflected bounce of light and suffer from systematic errors due to multipath interference. We re-purpose the conventional time of flight device for a new goal: to recover per-pixel sparse time profiles expressed as a sequence of impulses. With this modification, we show that we can not only address multipath interference but also enable new applications such as recovering depth of near-transparent surfaces, looking through diffusers and creating time-profile movies of sweeping light. Our key idea is to formulate the forward amplitude modulated light propagation as a convolution with custom codes, record samples by introducing a simple sequence of electronic time delays, and perform sparse deconvolution to recover sequences of Diracs that correspond to multipath returns. Applications to computer vision include ranging of near-transparent objects and subsurface imaging through diffusers. Our low cost prototype may lead to new insights regarding forward and inverse problems in light transport. Achuta Kadambi, Refael Whyte, Ayush Bhandari, Lee V. Streeter, Christopher Barsi, Adrian A. Dorrington, Ramesh Raskar |
ACM Trans. Graph. | 7 |
| 2013 | Focus 3D: Compressive accommodation displayabstractWe present a glasses-free 3D display design with the potential to provide viewers with nearly correct accommodative depth cues, as well as motion parallax and binocular cues. Building on multilayer attenuator and directional backlight architectures, the proposed design achieves the high angular resolution needed for accommodation by placing spatial light modulators about a large lens: one conjugate to the viewer's eye, and one or more near the plane of the lens. Nonnegative tensor factorization is used to compress a high angular resolution light field into a set of masks that can be displayed on a pair of commodity LCD panels. By constraining the tensor factorization to preserve only those light rays seen by the viewer, we effectively steer narrow high-resolution viewing cones into the user's eyes, allowing binocular disparity, motion parallax, and the potential for nearly correct accommodation over a wide field of view. We verify the design experimentally by focusing a camera at different depths about a prototype display, establish formal upper bounds on the design's accommodation range and diffraction-limited performance, and discuss practical limitations that must be overcome to allow the device to be used with human observers. Andrew Maimone, Gordon Wetzstein, Matthew Hirsch, Douglas Lanman, Ramesh Raskar, Henry Fuchs |
ACM Trans. Graph. | 5 |
| 2013 | Compressive light field photography using overcomplete dictionaries and optimized projectionsabstractLight field photography has gained a significant research interest in the last two decades; today, commercial light field cameras are widely available. Nevertheless, most existing acquisition approaches either multiplex a low-resolution light field into a single 2D sensor image or require multiple photographs to be taken for acquiring a high-resolution light field. We propose a compressive light field camera architecture that allows for higher-resolution light fields to be recovered than previously possible from a single image. The proposed architecture comprises three key components: light field atoms as a sparse representation of natural light fields, an optical design that allows for capturing optimized 2D light field projections, and robust sparse reconstruction methods to recover a 4D light field from a single coded 2D projection. In addition, we demonstrate a variety of other applications for light field atoms and sparse coding, including 4D light field compression and denoising. Kshitij Marwah, Gordon Wetzstein, Yosuke Bando, Ramesh Raskar |
ACM Trans. Graph. | 4 |
| 2013 | Femto-photography: capturing and visualizing the propagation of lightabstractWe present femto-photography , a novel imaging technique to capture and visualize the propagation of light. With an effective exposure time of 1.85 picoseconds (ps) per frame, we reconstruct movies of ultrafast events at an equivalent resolution of about one half trillion frames per second. Because cameras with this shutter speed do not exist, we re-purpose modern imaging hardware to record an ensemble average of repeatable events that are synchronized to a streak sensor, in which the time of arrival of light from the scene is coded in one of the sensor's spatial dimensions. We introduce reconstruction methods that allow us to visualize the propagation of femtosecond light pulses through macroscopic scenes; at such fast resolution, we must consider the notion of time-unwarping between the camera's and the world's space-time coordinate systems to take into account effects associated with the finite speed of light. We apply our femto-photography technique to visualizations of very different scenes, which allow us to observe the rich dynamics of time-resolved light transport effects, including scattering, specular reflections, diffuse interreflections, diffraction, caustics, and subsurface scattering. Our work has potential applications in artistic, educational, and scientific visualizations; industrial imaging to analyze material properties; and medical imaging to reconstruct subsurface elements. In addition, our time-resolved technique may motivate new forms of computational photography. Andreas Velten, Di Wu 0006, Adrián Jarabo, Belén Masiá, Christopher Barsi, Chinmaya Joshi, Everett Lawson, Moungi Bawendi, Diego Gutierrez, Ramesh Raskar |
ACM Trans. Graph. | 10 |
| 2013 | Using Patterns to Encode Color Information for DichromatsabstractColor is one of the most common ways to convey information in visualization applications. Color vision deficiency (CVD) affects approximately 200 million individuals worldwide and considerably degrades their performance in understanding such contents by creating red-green or blue-yellow ambiguities. While several content-specific methods have been proposed to resolve these ambiguities, they cannot achieve this effectively in many situations for contents with a large variety of colors. More importantly, they cannot facilitate color identification. We propose a technique for using patterns to encode color information for individuals with CVD, in particular for dichromats. We present the first content-independent method to overlay patterns on colored visualization contents that not only minimizes ambiguities but also allows color identification. Further, since overlaying patterns does not compromise the underlying original colors, it does not hamper the perception of normal trichromats. We validated our method with two user studies: one including 11 subjects with CVD and 19 normal trichromats, and focused on images that use colors to represent multiple categories; and another one including 16 subjects with CVD and 22 normal trichromats, which considered a broader set of images. Our results show that overlaying patterns significantly improves the performance of dichromats in several color-based visualization tasks, making their performance almost similar to normal trichromats'. More interestingly, the patterns augment color information in a positive manner, allowing normal trichromats to perform with greater accuracy. Behzad Sajadi, Aditi Majumder, Manuel Menezes de Oliveira Neto, Rosália G. Schneider, Ramesh Raskar |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2012 | Decomposing global light transport using time of flight imagingabstractGlobal light transport is composed of direct and indirect components. In this paper, we take the first steps toward analyzing light transport using high temporal resolution information via time of flight (ToF) images. The time profile at each pixel encodes complex interactions between the incident light and the scene geometry with spatially-varying material properties. We exploit the time profile to decompose light transport into its constituent direct, subsurface scattering, and interreflection components. We show that the time profile is well modelled using a Gaussian function for the direct and interreflection components, and a decaying exponential function for the subsurface scattering component. We use our direct, subsurface scattering, and interreflection separation algorithm for four computer vision applications: recovering projective depth maps, identifying subsurface scattering objects, measuring parameters of analytical subsurface scattering models, and performing edge detection using ToF images. Di Wu 0006, Matthew O'Toole, Andreas Velten, Amit K. Agrawal, Ramesh Raskar |
CVPR | 5 |
| 2012 | Frequency Analysis of Transient Light Transport with Applications in Bare Sensor Imaging
Di Wu 0006, Gordon Wetzstein, Christopher Barsi, Thomas Willwacher, Matthew O'Toole, Nikhil Naik 0003, Qionghai Dai, Kiriakos N. Kutulakos, Ramesh Raskar |
ECCV (1) | 9 |
| 2012 | Multimodal motion guidance: techniques for adaptive and dynamic feedbackabstractThe ability to guide human motion through automatically generated feedback has significant potential for applications in areas, such as motor learning, human-computer interaction, telepresence, and augmented reality. Christian Schönauer, Kenichiro Fukushi, Alex Olwal, Hannes Kaufmann, Ramesh Raskar |
ICMI | 5 |
| 2012 | VRCodes: Unobtrusive and active visual codes for interaction by exploiting rolling shutterabstractWe show a new visible tagging solution for active displays which allows a rolling-shutter camera to detect active tags from a relatively large distance in a robust manner. Current planar markers are visually obtrusive for the human viewer. In order for them to be read from afar and embed more information, they must be shown larger thus occupying valuable physical space on the design. We present a new active visual tag which utilizes all dimensions of color, time and space while remaining unobtrusive to the human eye and decodable using a 15fps rolling-shutter camera. The design exploits the flicker fusion-frequency threshold of the human visual system, which due to the effect of metamerism, can not resolve metamer pairs alternating beyond 120Hz. Yet, concurrently, it is decodable using a 15fps rolling-shutter camera due to the effective line-scan speed of 15×400 lines per second. We show an off-the-shelf rolling-shutter camera can resolve the metamers flickering on a television from a distance over 4 meters. We use intelligent binary coding to encode digital positioning and show potential applications such as large screen interaction. We analyze the use of codes for locking and tracking encoded targets. We also analyze the constraints and performance of the sampling system, and discuss several plausible application scenarios. Grace Woo, Andy Lippman, Ramesh Raskar |
ISMAR | 3 |
| 2012 | Reflectance model for diffractionabstractWe present a novel method of simulating wave effects in graphics using ray-based renderers with a new function: the Wave BSDF (Bidirectional Scattering Distribution Function). Reflections from neighboring surface patches represented by local BSDFs are mutually independent. However, in many surfaces with wavelength-scale microstructures, interference and diffraction requires a joint analysis of reflected wavefronts from neighboring patches. We demonstrate a simple method to compute the BSDF for the entire microstructure, which can be used independently for each patch. This allows us to use traditional ray-based rendering pipelines to synthesize wave effects. We exploit the Wigner Distribution Function (WDF) to create transmissive, reflective, and emissive BSDFs for various diffraction phenomena in a physically accurate way. In contrast to previous methods for computing interference, we circumvent the need to explicitly keep track of the phase of the wave by using BSDFs that include positive as well as negative coefficients. We describe and compare the theory in relation to well-understood concepts in rendering and demonstrate a straightforward implementation. In conjunction with standard raytracers, such as PBRT, we demonstrate wave effects for a range of scenarios such as multibounce diffraction materials, holograms, and reflection of high-frequency surfaces. Tom Cuypers, Tom Haber, Philippe Bekaert, Se Baek Oh, Ramesh Raskar |
ACM Trans. Graph. | 5 |
| 2012 | Correcting for optical aberrations using multilayer displaysabstractOptical aberrations of the human eye are currently corrected using eyeglasses, contact lenses, or surgery. We describe a fourth option: modifying the composition of displayed content such that the perceived image appears in focus, after passing through an eye with known optical defects. Prior approaches synthesize pre-filtered images by deconvolving the content by the point spread function of the aberrated eye. Such methods have not led to practical applications, due to severely reduced contrast and ringing artifacts. We address these limitations by introducing multilayer pre-filtering, implemented using stacks of semi-transparent, light-emitting layers. By optimizing the layer positions and the partition of spatial frequencies between layers, contrast is improved and ringing artifacts are eliminated. We assess design constraints for multilayer displays; autostereoscopic light field displays are identified as a preferred, thin form factor architecture, allowing synthetic layers to be displaced in response to viewer movement and refractive errors. We assess the benefits of multilayer pre-filtering versus prior light field pre-distortion methods, showing pre-filtering works within the constraints of current display resolutions. We conclude by analyzing benefits and limitations using a prototype multilayer LCD. Fu-Chung Huang, Douglas Lanman, Brian A. Barsky, Ramesh Raskar |
ACM Trans. Graph. | 4 |
| 2012 | Primal-dual coding to probe light transportabstractWe present primal-dual coding , a photography technique that enables direct fine-grain control over which light paths contribute to a photo. We achieve this by projecting a sequence of patterns onto the scene while the sensor is exposed to light. At the same time, a second sequence of patterns, derived from the first and applied in lockstep, modulates the light received at individual sensor pixels. We show that photography in this regime is equivalent to a matrix probing operation in which the elements of the scene's transport matrix are individually re-scaled and then mapped to the photo. This makes it possible to directly acquire photos in which specific light transport paths have been blocked, attenuated or enhanced. We show captured photos for several scenes with challenging light transport effects, including specular inter-reflections, caustics, diffuse inter-reflections and volumetric scattering. A key feature of primal-dual coding is that it operates almost exclusively in the optical domain: our results consist of directly-acquired, unprocessed RAW photos or differences between them. Matthew O'Toole, Ramesh Raskar, Kiriakos N. Kutulakos |
ACM Trans. Graph. | 2 |
| 2012 | Tailored displays to compensate for visual aberrationsabstractWe introduce tailored displays that enhance visual acuity by decomposing virtual objects and placing the resulting anisotropic pieces into the subject's focal range. The goal is to free the viewer from needing wearable optical corrections when looking at displays. Our tailoring process uses aberration and scattering maps to account for refractive errors and cataracts. It splits an object's light field into multiple instances that are each in-focus for a given eye sub-aperture. Their integration onto the retina leads to a quality improvement of perceived images when observing the display with naked eyes. The use of multiple depths to render each point of focus on the retina creates multi-focus, multi-depth displays. User evaluations and validation with modified camera optics are performed. We propose tailored displays for daily tasks where using eyeglasses are unfeasible or inconvenient (e.g., on head-mounted displays, e-readers, as well as for games); when a multi-focus function is required but undoable (e.g., driving for farsighted individuals, checking a portable device while doing physical activities); or for correcting the visual distortions produced by high-order aberrations that eyeglasses are not able to. Vitor F. Pamplona, Manuel Menezes de Oliveira Neto, Daniel G. Aliaga, Ramesh Raskar |
ACM Trans. Graph. | 4 |
| 2012 | Tensor displays: compressive light field synthesis using multilayer displays with directional backlightingabstractWe introduce tensor displays: a family of compressive light field displays comprising all architectures employing a stack of time-multiplexed, light-attenuating layers illuminated by uniform or directional backlighting (i.e., any low-resolution light field emitter). We show that the light field emitted by an N -layer, M -frame tensor display can be represented by an N th -order, rank- M tensor. Using this representation we introduce a unified optimization framework, based on nonnegative tensor factorization (NTF), encompassing all tensor display architectures. This framework is the first to allow joint multilayer, multiframe light field decompositions, significantly reducing artifacts observed with prior multilayer-only and multiframe-only decompositions; it is also the first optimization method for designs combining multiple layers with directional backlighting. We verify the benefits and limitations of tensor displays by constructing a prototype using modified LCD panels and a custom integral imaging backlight. Our efficient, GPU-based NTF implementation enables interactive applications. Through simulations and experiments we show that tensor displays reveal practical architectures with greater depths of field, wider fields of view, and thinner form factors, compared to prior automultiscopic displays. Gordon Wetzstein, Douglas Lanman, Matthew Hirsch, Ramesh Raskar |
ACM Trans. Graph. | 4 |
| 2011 | Kinected conference: augmenting video imaging with calibrated depth and audioabstractThe proliferation of broadband and high-speed Internet access has, in general, democratized the ability to commonly engage in videoconference. However, current video systems do not meet their full potential, as they are restricted to a simple display of unintelligent 2D pixels. In this paper we present a system for enhancing distance-based communication by augmenting the traditional video conferencing system with additional attributes beyond two-dimensional video. We explore how expanding a system's understanding of spatially calibrated depth and audio alongside a live video stream can generate semantically rich three-dimensional pixels containing information regarding their material properties and location. We discuss specific scenarios that explore features such as synthetic refocusing, gesture activated privacy, and spatiotemporal graphic augmentation. Anthony DeVincenzi, Lining Yao, Hiroshi Ishii 0001, Ramesh Raskar |
CSCW | 4 |
| 2011 | Estimating Motion and size of moving non-line-of-sight objects in cluttered environmentsabstractWe present a technique for motion and size estimation of non-line-of-sight (NLOS) moving objects in cluttered environments using a time of flight camera and multipath analysis. We exploit relative times of arrival after reflection from a grid of points on a diffuse surface and create a virtual phased-array. By subtracting space-time impulse responses for successive frames, we separate responses of NLOS moving objects from those resulting from the cluttered environment. After reconstructing the line-of-sight scene geometry, we analyze the space of wavefronts using the phased array and solve a constrained least squares problem to recover the NLOS target location. Importantly, we can recover target's motion vector even in presence of uncalibrated time and pose bias common in time of flight systems. In addition, we compute the upper bound on the size of the target by backprojecting the extremas of the time profiles. Ability to track targets inside rooms despite opaque occluders and multipath responses has numerous applications in search and rescue, medicine and defense. We show centimeter accurate results by making appropriate modifications to a time of flight system. Rohit Pandharkar, Andreas Velten, Andrew Bardagjy, Everett Lawson, Moungi Bawendi, Ramesh Raskar |
CVPR | 6 |
| 2011 | Validity of Wigner Distribution Function for ray-based imagingabstractIn this work we provide an introduction to the Wigner Distribution Function (WDF) using geometric optics principles. The WDF provides a useful model of wave-fields, allowing simulation of diffraction and interference effects. We attempt to explain these Fourier optics concepts to computational photography researchers by clarifying the relationship between the WDF and position-angle representations. We demonstrate how the WDF can be used to simulate diffraction effects using a light field representation and discuss its validity in the near-field, far-field, and under the paraxial approximation. Finally, we demonstrate that although the WDF representation contains negative values, any projection always yields a non-negative intensity value. Tom Cuypers, Roarke Horstmeyer, Se Baek Oh, Philippe Bekaert, Ramesh Raskar |
ICCP | 5 |
| 2011 | Hand-held Schlieren Photography with Light Field probesabstractWe introduce a new approach to capturing refraction in transparent media, which we call Light Field Background Oriented Schlieren Photography (LFBOS). By optically coding the locations and directions of light rays emerging from a light field probe, we can capture changes of the refractive index field between the probe and a camera or an observer. Rather than using complicated and expensive optical setups as in traditional Schlieren photography we employ commodity hardware; our prototype consists of a camera and a lenslet array. By carefully encoding the color and intensity variations of a 4D probe instead of a diffuse 2D background, we avoid expensive computational processing of the captured data, which is necessary for Background Oriented Schlieren imaging (BOS). We analyze the benefits and limitations of our approach and discuss application scenarios. Gordon Wetzstein, Ramesh Raskar, Wolfgang Heidrich |
ICCP | 2 |
| 2011 | Refractive shape from light field distortionabstractAcquiring transparent, refractive objects is challenging as these kinds of objects can only be observed by analyzing the distortion of reference background patterns. We present a new, single image approach to reconstructing thin transparent surfaces, such as thin solids or surfaces of fluids. Our method is based on observing the distortion of light field background illumination. Light field probes have the potential to encode up to four dimensions in varying colors and intensities: spatial and angular variation on the probe surface; commonly employed reference patterns are only two-dimensional by coding either position or angle on the probe. We show that the additional information can be used to reconstruct refractive surface normals and a sparse set of control points from a single photograph. Gordon Wetzstein, David Roodnick, Wolfgang Heidrich, Ramesh Raskar |
ICCV | 4 |
| 2011 | SpeckleSense: fast, precise, low-cost and compact motion sensing using laser speckleabstractMotion sensing is of fundamental importance for user interfaces and input devices. In applications, where optical sensing is preferred, traditional camera-based approaches can be prohibitive due to limited resolution, low frame rates and the required computational power for image processing. We introduce a novel set of motion-sensing configurations based on laser speckle sensing that are particularly suitable for human-computer interaction. The underlying principles allow these configurations to be fast, precise, extremely compact and low cost. We provide an overview and design guidelines for laser speckle sensing for user interaction and introduce four general speckle projector/sensor configurations. We describe a set of prototypes and applications that demonstrate the versatility of our laser speckle sensing techniques. Jan Zizka, Alex Olwal, Ramesh Raskar |
UIST | 3 |
| 2011 | Dynamic Display of BRDFsabstractAbstract This paper deals with the challenge of physically displaying reflectance, i.e., the appearance of a surface and its variation with the observer position and the illuminating environment. This is commonly described by the bidirectional reflectance distribution function (BRDF). We provide a catalogue of criteria for the display of BRDFs, and sketch a few orthogonal approaches to solving the problem in an optically passive way. Our specific implementation is based on a liquid surface, on which we excite waves in order to achieve a varying degree of anisotropic roughness. The resulting probability density function of the surface normal is shown to follow a Gaussian distribution similar to most established BRDF models. Matthias B. Hullin, Hendrik P. A. Lensch, Ramesh Raskar, Hans-Peter Seidel, Ivo Ihrke |
Comput. Graph. Forum | 3 |
| 2011 | Looking Around the Corner using Ultrafast Transient Imaging
Ahmed Kirmani, Tyler Hutchison, James Davis 0001, Ramesh Raskar |
Int. J. Comput. Vis. | 4 |
| 2011 | Coded Strobing Photography: Compressive Sensing of High Speed Periodic VideosabstractWe show that, via temporal modulation, one can observe and capture a high-speed periodic video well beyond the abilities of a low-frame-rate camera. By strobing the exposure with unique sequences within the integration time of each frame, we take coded projections of dynamic events. From a sequence of such frames, we reconstruct a high-speed video of the high-frequency periodic process. Strobing is used in entertainment, medical imaging, and industrial inspection to generate lower beat frequencies. But this is limited to scenes with a detectable single dominant frequency and requires high-intensity lighting. In this paper, we address the problem of sub-Nyquist sampling of periodic signals and show designs to capture and reconstruct such signals. The key result is that for such signals, the Nyquist rate constraint can be imposed on the strobe rate rather than the sensor rate. The technique is based on intentional aliasing of the frequency components of the periodic signal while the reconstruction algorithm exploits recent advances in sparse representations and compressive sensing. We exploit the sparsity of periodic signals in the Fourier domain to develop reconstruction algorithms that are inspired by compressive sensing. Ashok Veeraraghavan, Dikpal Reddy, Ramesh Raskar |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | Highlighted depth-of-field photography: Shining light on focusabstractWe present a photographic method to enhance intensity differences between objects at varying distances from the focal plane. By combining a unique capture procedure with simple image processing techniques, the detected brightness of an object is decreased proportional to its degree of defocus. A camera-projector system casts distinct grid patterns onto a scene to generate a spatial distribution of point reflections. These point reflections relay a relative measure of defocus that is utilized in postprocessing to generate a highlighted DOF photograph. Trade-offs between three different projector-processing pairs are analyzed, and a model is developed to help describe a new intensity-dependent depth of field that is controlled by the pattern of illumination. Results are presented for a primary single snapshot design as well as a scanning method and a comparison method. As an application, automatic matting results are presented. Roarke Horstmeyer, Ig-Jae Kim, Ramesh Raskar |
ACM Trans. Graph. | 4 |
| 2011 | Polarization fields: dynamic light field display using multi-layer LCDsabstractWe introduce polarization field displays as an optically-efficient design for dynamic light field display using multi-layered LCDs. Such displays consist of a stacked set of liquid crystal panels with a single pair of crossed linear polarizers. Each layer is modeled as a spatially-controllable polarization rotator, as opposed to a conventional spatial light modulator that directly attenuates light. Color display is achieved using field sequential color illumination with monochromatic LCDs, mitigating severe attenuation and moiré occurring with layered color filter arrays. We demonstrate such displays can be controlled, at interactive refresh rates, by adopting the SART algorithm to tomographically solve for the optimal spatially-varying polarization state rotations applied by each layer. We validate our design by constructing a prototype using modified off-the-shelf panels. We demonstrate interactive display using a GPU-based SART implementation supporting both polarization-based and attenuation-based architectures. Experiments characterize the accuracy of our image formation model, verifying polarization field displays achieve increased brightness, higher resolution, and extended depth of field, as compared to existing automultiscopic display methods for dual-layer and multi-layer LCDs. Douglas Lanman, Gordon Wetzstein, Matthew Hirsch, Wolfgang Heidrich, Ramesh Raskar |
ACM Trans. Graph. | 5 |
| 2011 | Single view reflectance capture using multiplexed scattering and time-of-flight imagingabstractThis paper introduces the concept of time-of-flight reflectance estimation, and demonstrates a new technique that allows a camera to rapidly acquire reflectance properties of objects from a single view-point, over relatively long distances and without encircling equipment. We measure material properties by indirectly illuminating an object by a laser source, and observing its reflected light indirectly using a time-of-flight camera. The configuration collectively acquires dense angular, but low spatial sampling, within a limited solid angle range - all from a single viewpoint. Our ultra-fast imaging approach captures space-time "streak images" that can separate out different bounces of light based on path length. Entanglements arise in the streak images mixing signals from multiple paths if they have the same total path length. We show how reflectances can be recovered by solving for a linear system of equations and assuming parametric material models; fitting to lower dimensional reflectance models enables us to disentangle measurements. We demonstrate proof-of-concept results of parametric reflectance models for homogeneous and discretized heterogeneous patches, both using simulation and experimental hardware. As compared to lengthy or highly calibrated BRDF acquisition techniques, we demonstrate a device that can rapidly, on the order of seconds, capture meaningful reflectance information. We expect hardware advances to improve the portability and speed of this device. Nikhil Naik 0003, Andreas Velten, Ramesh Raskar, Kavita Bala |
ACM Trans. Graph. | 4 |
| 2011 | CATRA: interactive measuring and modeling of cataractsabstractWe introduce an interactive method to assess cataracts in the human eye by crafting an optical solution that measures the perceptual impact of forward scattering on the foveal region. Current solutions rely on highly-trained clinicians to check the back scattering in the crystallin lens and test their predictions on visual acuity tests. Close-range parallax barriers create collimated beams of light to scan through sub-apertures, scattering light as it strikes a cataract. User feedback generates maps for opacity, attenuation, contrast and sub-aperture point-spread functions. The goal is to allow a general audience to operate a portable high-contrast light-field display to gain a meaningful understanding of their own visual conditions. User evaluations and validation with modified camera optics are performed. Compiled data is used to reconstruct the individual's cataract-affected view, offering a novel approach for capturing information for screening, diagnostic, and clinical analysis. Vitor F. Pamplona, Erick Baptista Passos, Jan Zizka, Manuel Menezes de Oliveira Neto, Everett Lawson, Esteban Walter Gonzalez Clua, Ramesh Raskar |
ACM Trans. Graph. | 7 |
| 2011 | Switchable primaries using shiftable layers of color filter arraysabstractWe present a camera with switchable primaries using shiftable layers of color filter arrays (CFAs). By layering a pair of CMY CFAs in this novel manner we can switch between multiple sets of color primaries (namely RGB, CMY and RGBCY) in the same camera. In contrast to fixed color primaries (e.g. RGB or CMY), which cannot provide optimal image quality for all scene conditions, our camera with switchable primaries provides optimal color fidelity and signal to noise ratio for multiple scene conditions. Next, we show that the same concept can be used to layer two RGB CFAs to design a camera with switchable low dynamic range (LDR) and high dynamic range (HDR) modes. Further, we show that such layering can be generalized as a constrained satisfaction problem (CSP) allowing to constrain a large number of parameters (e.g. different operational modes, amount and direction of the shifts, placement of the primaries in the CFA) to provide an optimal solution. We investigate practical design options for shiftable layering of the CFAs. We demonstrate these by building prototype cameras for both switchable primaries and switchable LDR/HDR modes. To the best of our knowledge, we present, for the first time, the concept of shiftable layers of CFAs that provides a new degree of freedom in photography where multiple operational modes are available to the user in a single camera for optimizing the picture quality based on the nature of the scene geometry, color and illumination. Behzad Sajadi, Aditi Majumder, Kazuhiro Hiwada, Atsuto Maki, Ramesh Raskar |
ACM Trans. Graph. | 5 |
| 2011 | Layered 3D: tomographic image synthesis for attenuation-based light field and high dynamic range displaysabstractWe develop tomographic techniques for image synthesis on displays composed of compact volumes of light-attenuating material. Such volumetric attenuators recreate a 4D light field or high-contrast 2D image when illuminated by a uniform backlight. Since arbitrary oblique views may be inconsistent with any single attenuator, iterative tomographic reconstruction minimizes the difference between the emitted and target light fields, subject to physical constraints on attenuation. As multi-layer generalizations of conventional parallax barriers, such displays are shown, both by theory and experiment, to exceed the performance of existing dual-layer architectures. For 3D display, spatial resolution, depth of field, and brightness are increased, compared to parallax barriers. For a plane at a fixed depth, our optimization also allows optimal construction of high dynamic range displays, confirming existing heuristics and providing the first extension to multiple, disjoint layers. We conclude by demonstrating the benefits and limitations of attenuation-based light field displays using an inexpensive fabrication method: separating multiple printed transparencies with acrylic sheets. Gordon Wetzstein, Douglas Lanman, Wolfgang Heidrich, Ramesh Raskar |
ACM Trans. Graph. | 4 |
| 2010 | Hemispherical Confocal Imaging Using Turtleback Reflector
Yasuhiro Mukaigawa, Seiichi Tagawa, Ramesh Raskar, Yasuyuki Matsushita, Yasushi Yagi |
ACCV (1) | 4 |
| 2010 | Analysis of light transport in scattering mediaabstractWe propose a new method to analyze light transport in homogeneous scattering media. The incident light undergoes multiple bounces in translucent objects, and produces a complex light field. Our method analyzes the light transport in two steps. First, single and multiple scattering are separated by projecting high-frequency stripe patterns. Then, multiple scattering is decomposed into each bounce component based on the light transport equation. The light field for each bounce is recursively estimated. Experimental results show that light transport in scattering media can be decomposed and visualized for each bounce. Yasuhiro Mukaigawa, Yasushi Yagi, Ramesh Raskar |
CVPR | 3 |
| 2010 | Descattering Transmission via Angular Filtering
Douglas Lanman, Yasuhiro Mukaigawa, Ramesh Raskar |
ECCV (1) | 4 |
| 2010 | Image sensing under unfavorable photographic conditions with a group of wireless image sensorsabstractWe investigate the problem of image sensing under unfavorable photographic conditions in a wireless image sensor network. In the scenes with deflective and/or reflective medium such as fogs, mirrors, glasses, degraded images are captured by those image sensors. Such degraded images often lack perceptual vividness and they offer a poor visibility of the scene contents. Notably, computation-intensive method to recover a better image based on single image [2] may not be applicable for wireless image sensors due to the limited computation capacities and the limited power resources (batteries) typically equipped at those wireless image sensors. In this paper, we propose a framework to recover better images under unfavorable photographic conditions in a wireless image sensor network, where an efficient decision fusion approach based on the reinforcement learning technique to infer the presence of an unfavorable photographic condition is achieved among image sensors on the fly and a subsequent light-weighted computation method based on multiple images is employed to recover better images. The preliminary results show the effectiveness of the proposed framework. Fulu Li, James Barabas, Ankit Mohan, Ramesh Raskar |
IPSN | 4 |
| 2010 | Second Skin: Motion capture with actuated feedback for motor learningabstractSecond Skin aims to combine three-dimensional motion tracking with real-time tactile feedback for the purpose of improving a user's motor-learning ability. Body and limb movements are tracked in 3D as a user performs an action, and the user is given automatic, real-time tactile feedback to aid in the correction of movement and position errors. The key areas of the system involve the development of a robust optical motion tracking system that can be used in a wide range of environments and lighting conditions, and an effective tactile feedback method capable of delivering information to the user to indicate how to correct errors in body and limb position. Since a number of components of the motion tracking and tactile feedback systems must be bound to the user's body, another important goal is the design of a lightweight and minimally inhibitive wearable suit to contain all of these elements. Dennis R. Miaw, Ramesh Raskar |
VR | 2 |
| 2010 | Real-time temporal shaping of high-speed video streams
Martin Fuchs 0001, Tongbo Chen, Oliver Wang, Ramesh Raskar, Hans-Peter Seidel, Hendrik P. A. Lensch |
Comput. Graph. | 4 |
| 2010 | Reinterpretable Imager: Towards Variable Post-Capture Space, Angle and Time Resolution in PhotographyabstractAbstract We describe a novel multiplexing approach to achieve tradeoffs in space, angle and time resolution in photography. We explore the problem of mapping useful subsets of time‐varying 4D lightfields in a single snapshot. Our design is based on using a dynamic mask in the aperture and a static mask close to the sensor. The key idea is to exploit scene‐specific redundancy along spatial, angular and temporal dimensions and to provide a programmable or variable resolution tradeoff among these dimensions. This allows a user to reinterpret the single captured photo as either a high spatial resolution image, a refocusable image stack or a video for different parts of the scene in post‐processing. A lightfield camera or a video camera forces a‐priori choice in space‐angle‐time resolution. We demonstrate a single prototype which provides flexible post‐capture abilities not possible using either a single‐shot lightfield camera or a multi‐frame video camera. We show several novel results including digital refocusing on objects moving in depth and capturing multiple facial expressions in a single photo. Amit K. Agrawal, Ashok Veeraraghavan, Ramesh Raskar |
Comput. Graph. Forum | 3 |
| 2010 | Rendering Wave Effects with Augmented Light FieldabstractAbstract Ray–based representations can model complex light transport but are limited in modeling diffraction effects that require the simulation of wavefront propagation. This paper provides a new paradigm that has the simplicity of light path tracing and yet provides an accurate characterization of both Fresnel and Fraunhofer diffraction. We introduce the concept of a light field transformer at the interface of transmissive occluders. This generates mathematically sound, virtual, and possibly negative‐valued light sources after the occluder. From a rendering perspective the only simple change is that radiance can be temporarily negative. We demonstrate the correctness of our approach both analytically, as well by comparing values with standard experiments in physics such as the Young's double slit. Our implementation is a shader program in OpenGL that can generate wave effects on arbitrary surfaces. Se Baek Oh, Sriram Kashyap, Rohit Garg, Sharat Chandran, Ramesh Raskar |
Comput. Graph. Forum | 5 |
| 2010 | Content-adaptive parallax barriers: optimizing dual-layer 3D displays using low-rank light field factorizationabstractWe optimize automultiscopic displays built by stacking a pair of modified LCD panels. To date, such dual-stacked LCDs have used heuristic parallax barriers for view-dependent imagery: the front LCD shows a fixed array of slits or pinholes, independent of the multi-view content. While prior works adapt the spacing between slits or pinholes, depending on viewer position, we show both layers can also be adapted to the multi-view content, increasing brightness and refresh rate. Unlike conventional barriers, both masks are allowed to exhibit non-binary opacities. It is shown that any 4D light field emitted by a dual-stacked LCD is the tensor product of two 2D masks. Thus, any pair of 1D masks only achieves a rank-1 approximation of a 2D light field. Temporal multiplexing of masks is shown to achieve higher-rank approximations. Non-negative matrix factorization (NMF) minimizes the weighted Euclidean distance between a target light field and that emitted by the display. Simulations and experiments characterize the resultingcontent-adaptive parallax barriersfor low-rank light field approximation. Douglas Lanman, Matthew Hirsch, Yunhee Kim, Ramesh Raskar |
ACM Trans. Graph. | 4 |
| 2010 | NETRA: interactive display for estimating refractive errors and focal rangeabstractWe introduce an interactive, portable, and inexpensive solution for estimating refractive errors in the human eye. While expensive optical devices for automatic estimation of refractive correction exist, our goal is to greatly simplify the mechanism by putting the human subject in the loop. Our solution is based on a high-resolution programmable display and combines inexpensive optical elements, interactive GUI, and computational reconstruction. The key idea is to interface a lenticular view-dependent display with the human eye in close range - a few millimeters apart. Via this platform, we create a new range of interactivity that is extremely sensitive to parameters of the human eye, like refractive errors, focal range, focusing speed, lens opacity, etc. We propose several simple optical setups, verify their accuracy, precision, and validate them in a user study. Vitor F. Pamplona, Ankit Mohan, Manuel Menezes de Oliveira Neto, Ramesh Raskar |
ACM Trans. Graph. | 4 |
| 2010 | Axial-cones: modeling spherical catadioptric cameras for wide-angle light field renderingabstractCatadioptric imaging systems are commonly used for wide-angle imaging, but lead to multi-perspective images which do not allow algorithms designed for perspective cameras to be used. Efficient use of such systems requires accurate geometric ray modeling as well as fast algorithms. We present accurate geometric modeling of the multi-perspective photo captured with a spherical catadioptric imaging system usingaxial-cone cameras:multiple perspective cameras lying on an axis each with a different viewpoint and a different cone of rays. This modeling avoids geometric approximations and allows several algorithms developed for perspective cameras to be applied to multi-perspective catadioptric cameras. We demonstrate axial-cone modeling in the context of rendering wide-angle light fields, captured using a spherical mirror array. We present several applications such as spherical distortion correction, digital refocusing for artistic depth of field effects in wide-angle scenes, and wide-angle dense depth estimation. Our GPU implementation using axial-cone modeling achieves up to three orders of magnitude speed up over ray tracing for these applications. Yuichi Taguchi, Amit K. Agrawal, Ashok Veeraraghavan, Srikumar Ramalingam, Ramesh Raskar |
ACM Trans. Graph. | 5 |
| 2009 | Optimal single image capture for motion deblurringabstractDeblurring images of moving objects captured from a traditional camera is an ill-posed problem due to the loss of high spatial frequencies in the captured images. Techniques have attempted to engineer the motion point spread function (PSF) by either making it invertible using coded exposure, or invariant to motion by moving the camera in a specific fashion. We address the problem of optimal single image capture strategy for best deblurring performance. We formulate the problem of optimal capture as maximizing the signal to noise ratio (SNR) of the deconvolved image given a scene light level. As the exposure time increases, the sensor integrates more light, thereby increasing the SNR of the captured signal. However, for moving objects, larger exposure time also results in more blur and hence more deconvolution noise. We compare the following three single image capture strategies: (a) traditional camera, (b) coded exposure camera, and (c) motion invariant photography, as well as the best exposure time for capture by analyzing the rate of increase of deconvolution noise with exposure time. We analyze which strategy is optimal for known/unknown motion direction and speed and investigate how the performance degrades for other cases. We present real experimental results by simulating the above capture strategies using a high speed video camera. Amit K. Agrawal, Ramesh Raskar |
CVPR | 2 |
| 2009 | 3D pose estimation and segmentation using specular cuesabstractWe present a system for fast model-based segmentation and 3D pose estimation of specular objects using appearance based specular features. We use observed (a) specular reflection and (b) specular flow as cues, which are matched against similar cues generated from a CAD model of the object in various poses. We avoid estimating 3D geometry or depths, which is difficult and unreliable for specular scenes. In the first method, the environment map of the scene is utilized to generate a database containing synthesized specular reflections of the object for densely sampled 3D poses. This database is compared with captured images of the scene at run time to locate and estimate the 3D pose of the object. In the second method, specular flows are generated for dense 3D poses as illumination invariant features and are matched to the specular flow of the scene. We incorporate several practical heuristics such as use of saturated/highlight pixels for fast matching and normal selection to minimize the effects of inter-reflections and cluttered backgrounds. Despite its simplicity, our approach is effective in scenes with multiple specular objects, partial occlusions, inter-reflections, cluttered backgrounds and changes in ambient illumination. Experimental results demonstrate the effectiveness of our method for various synthetic and real objects. Ju Yong Chang, Ramesh Raskar, Amit K. Agrawal |
CVPR | 2 |
| 2009 | A projector-camera setup for geometry-invariant frequency demultiplexingabstractConsider a projector-camera setup where a sinusoidal pattern is projected onto the scene, and an image of the objects imprinted with the pattern is captured by the camera. In this configuration, the local frequency of the sinusoidal pattern as seen by the camera is a function of both the frequency of the projected sinusoid and the local geometry of objects in the scene. We observe that, by strategically placing the projector and the camera in canonical configuration and projecting sinusoidal patterns aligned with the epipolar lines, the frequency of the sinusoids seen in the image becomes invariant to the local object geometry. This property allows us to design systems composed of a camera and multiple projectors, which can be used to capture a single image of a scene illuminated by all projectors at the same time, and then demultiplex the frequencies generated by each individual projector separately. We show how imaging systems like those can be used to segment, from a single image, the shadows cast by each individual projector - an application that we call coded shadow photography. The method is useful to extend the applicability of techniques that rely on the analysis of shadows cast by multiple light sources placed at different positions, as the individual shadows captured at distinct instants of time now can be obtained from a single shot, enabling the processing of dynamic scenes. Daniel A. Vaquero, Ramesh Raskar, Rogério Feris, Matthew Turk 0001 |
CVPR | 2 |
| 2009 | Looking around the corner using transient imagingabstractWe show that multi-path analysis using images from a timeof-flight (ToF) camera provides a tantalizing opportunity to infer about 3D geometry of not only visible but hidden parts of a scene. We provide a novel framework for reconstructing scene geometry from a single viewpoint using a camera that captures a 3D time-image I(x, y, t) for each pixel. We propose a framework that uses the time-image and transient reasoning to expose scene properties that may be beyond the reach of traditional computer vision. We corroborate our theory with free space hardware experiments using a femtosecond laser and an ultrafast photo detector array. The ability to compute the geometry of hidden elements, unobservable by both the camera and illumination source, will create a range of new computer vision opportunities. Ahmed Kirmani, Tyler Hutchison, James Davis 0001, Ramesh Raskar |
ICCV | 4 |
| 2009 | A Unified Calibration Method with a Parametric Approach for Wide-Field-of-View Multiprojector DisplaysabstractIn this paper, we describe techniques for supporting a wide-field-of-view multiprojector curved screen display system. Our main contribution is in achieving automatic geometric calibration and efficient rendering for seamless displays, which is effective even in the presence of panoramic surround screens with the multiview calibration method without polygonal representation of the display surface. We show several prototype systems that use a stereo camera for capturing and a new rendering method for quadric curved screens. Previous approaches have required a calibration camera at the sweet spot. Due to parameterized representation, however, our unified calibration method is independent of the orientation and field of view of the calibration camera. This method can simplify the tedious and complicated installation process as well as the maintenance of large multiprojector displays in planetariums, virtual reality systems, and other visualization venues. Masato Ogata, Hiroyuki Wada, Jeroen van Baar, Ramesh Raskar |
VR | 4 |
| 2009 | Invertible motion blur in videoabstractWe show that motion blur in successive video frames is invertible even if the point-spread function (PSF) due to motion smear in a single photo is non-invertible. Blurred photos exhibit nulls (zeros) in the frequency transform of the PSF, leading to an ill-posed deconvolution. Hardware solutions to avoid this require specialized devices such as the coded exposure camera or accelerating sensor motion. We employ ordinary video cameras and introduce the notion of null-filling along with joint-invertibility of multiple blur-functions. The key idea is to record the same object with varying PSFs, so that the nulls in the frequency component of one frame can be filled by other frames. The combined frequency transform becomes null-free, making deblurring well-posed. We achieve jointly-invertible blur simply by changing the exposure time of successive frames. We address the problem of automatic deblurring of objects moving with constant velocity by solving the four critical components: preservation of all spatial frequencies, segmentation of moving parts, motion estimation of moving parts, and non-degradation of the static parts of the scene. We demonstrate several challenging cases of object motion blur including textured backgrounds and partial occluders. Amit K. Agrawal, Ramesh Raskar |
ACM Trans. Graph. | 3 |
| 2009 | BiDi screen: a thin, depth-sensing LCD for 3D interaction using light fieldsabstractWe transform an LCD into a display that supports both 2D multi-touch and unencumbered 3D gestures. Our BiDirectional (BiDi) screen, capable of both image capture and display, is inspired by emerging LCDs that use embedded optical sensors to detect multiple points of contact. Our key contribution is to exploit the spatial light modulation capability of LCDs to allow lensless imaging without interfering with display functionality. We switch between a display mode showing traditional graphics and a capture mode in which the backlight is disabled and the LCD displays a pinhole array or an equivalent tiled-broadband code. A large-format image sensor is placed slightly behind the liquid crystal layer. Together, the image sensor and LCD form a mask-based light field camera, capturing an array of images equivalent to that produced by a camera array spanning the display surface. The recovered multi-view orthographic imagery is used to passively estimate the depth of scene points. Two motivating applications are described: a hybrid touch plus gesture interaction and a light-gun mode for interacting with external light-emitting widgets. We show a working prototype that simulates the image sensor with a camera and diffuser, allowing interaction up to 50 cm in front of a modified 20.1 inch LCD. Matthew Hirsch, Douglas Lanman, Henry Holtzman, Ramesh Raskar |
ACM Trans. Graph. | 4 |
| 2009 | Bokode: imperceptible visual tags for camera based interaction from a distanceabstractWe show a new camera based interaction solution where an ordinary camera can detect small optical tags from a relatively large distance. Current optical tags, such as barcodes, must be read within a short range and the codes occupy valuable physical space on products. We present a new low-cost optical design so that the tags can be shrunk to 3mm visible diameter, and unmodified ordinary cameras several meters away can be set up to decode the identity plus the relative distance and angle. The design exploits the bokeh effect of ordinary cameras lenses, which maps rays exiting from an out of focus scene point into a disk like blur on the camera sensor. This bokeh-code or Bokode is a barcode design with a simple lenslet over the pattern. We show that a code with 15 μm features can be read using an off-the-shelf camera from distances of up to 2 meters. We use intelligent binary coding to estimate the relative distance and angle to the camera, and show potential for applications in augmented reality and motion capture. We analyze the constraints and performance of the optical system, and discuss several plausible application scenarios. Ankit Mohan, Grace Woo, Shinsaku Hiura, Quinn Smithwick, Ramesh Raskar |
ACM Trans. Graph. | 5 |
| 2008 | Sensing increased image resolution using aperture masksabstractWe present a technique to construct increased-resolution images from multiple photos taken without moving the camera or the sensor. Like other super-resolution techniques, we capture and merge multiple images, but instead of moving the camera sensor by sub-pixel distances for each image, we change masks in the lens aperture and slightly defocus the lens. The resulting capture system is simpler, and tolerates modest mask registration errors well. We present a theoretical analysis of the camera and image merging method, show both simulated results and actual results from a crudely modified consumer camera, and compare its results to robust ‘blind’ methods that rely on uncontrolled camera displacements. Ankit Mohan, Xiang Huang 0006, Jack Tumblin, Ramesh Raskar |
CVPR | 4 |
| 2008 | Characterizing the shadow space of camera-light pairsabstractWe present a theoretical analysis for characterizing the shadows cast by a point light source given its relative position to the camera. In particular, we analyze the epipolar geometry of camera-light pairs, including unusual camera-light configurations such as light sources aligned with the camerapsilas optical axis as well as convenient arrangements such as lights placed in the camera plane. A mathematical characterization of the shadows is derived to determine the orientations and locations of depth discontinuities when projected onto the image plane that could potentially be associated with cast shadows. The resulting theory is applied to compute a lower bound on the number of lights needed to extract all depth discontinuities from a general scene using a multiflash camera. We also provide a characterization of which discontinuities are missed and which are correctly detected by the algorithm, and a foundation for choosing an optimal light placement. Experiments with depth edges computed using two-flash setups and a four-flash setup illustrate the theory, and an additional configuration with a flash at the camerapsilas center of projection is exploited as a solution for some degenerate cases. Daniel A. Vaquero, Rogério Feris, Matthew Turk 0001, Ramesh Raskar |
CVPR | 4 |
| 2008 | Non-refractive modulators for encoding and capturing scene appearance and depthabstractWe analyze the modulation of a light field via non-refracting attenuators. In the most general case, any desired modulation can be achieved with attenuators having four degrees of freedom in ray-space. We motivate the discussion with a universal 4D ray modulator (ray-filter) which can attenuate the intensity of each ray independently. We describe operation of such a fantasy ray-filter in the context of altering the 4D light field incident on a 2D camera sensor. Ray-filters are difficult to realize in practice but we can achieve reversible encoding for light field capture using patterned attenuating mask. Two mask-based designs are analyzed in this framework. The first design closely mimics the angle-dependent ray-sorting possible with the ray filter. The second design exploits frequency-domain modulation to achieve a more efficient encoding. We extend these designs for optimal sampling of light field by matching the modulation function to the specific shape of the band-limit frequency transform of light field. We also show how a hand-held version of an attenuator based light field camera can be built using a medium-format digital camera and an inexpensive mask. Ashok Veeraraghavan, Amit K. Agrawal, Ramesh Raskar, Ankit Mohan, Jack Tumblin |
CVPR | 3 |
| 2008 | Analysis on Probabilistic View Coverage for Image Sensing - A Geometric ApproachabstractIn this paper we study the probabilistic view coverage problem for image sensing in wireless sensor networks. The view coverage of an image sensor network determines the quality of the surveillance services that an image sensor network can provide. In this paper, we present an indepth analysis on probabilistic view coverage in an image sensor network, where omnidirectional image sensors are randomly dropped to a given field and the locations of the image sensors may not be immediately known. We intend to answer the following question: if we randomly drop a given number of image sensors into a targeted field, what is the probability that a given area of interest can be effectively imaged and view-covered. The key to our analytical approach is to cast the probabilistic view coverage problem in wireless image sensor networks as a geometric one and then use the geometric techniques to find the solution. The analysis in this paper provides probabilistic assurance of the view coverage that one can expect for random dropping of omnidirectional image sensors into a given field. Fulu Li, Ramesh Raskar, Andy Lippman |
IPCCC | 2 |
| 2008 | Capturing Images with Sparse Informational Pixels using Projected 3D TagsabstractIn this paper, we propose a novel imaging system that enables the capture of photos and videos with sparse informational pixels. Our system is based on the projection and detection of 3D optical tags. We use an infrared (IR) projector to project temporally-coded (blinking) dots onto selected points in a scene. These tags are invisible to the human eye, but appear as clearly visible time-varying codes to an IR photosensor. As a proof of concept, we have built a prototype camera system (consisting of co-located visible and IR sensors) to simultaneously capture visible and IR images. When a user takes an image of a tagged scene using such a camera system, all the scene tags that are visible from the system's viewpoint are detected. In addition, tags that lie in the field of view but are occluded, and ones that lie just outside the field of view, are also automatically generated for the image. Associated with each tagged pixel is its 3D location and the identity of the object that the tag falls on. Our system can interface with conventional image recognition methods for efficient scene authoring, enabling objects in an image to be robustly identified using cheap cameras, minimal computations, and no domain knowledge. We demonstrate several applications of our system, including, photo-browsing, e-commerce, augmented reality, and objection localization. Li Zhang 0003, Neesha Subramaniam, Robert Lin, Ramesh Raskar, Shree K. Nayar |
VR | 4 |
| 2008 | Agile Spectrum Imaging: Programmable Wavelength Modulation for Cameras and ProjectorsabstractAbstract We advocate the use of quickly‐adjustable, computer‐controlled color spectra in photography, lighting and displays. We present an optical relay system that allows mechanical or electronic color spectrum control and use it to modify a conventional camera and projector. We use a diffraction grating to disperse the rays into different colors, and introduce a mask (or LCD/DMD) in the optical path to modulate the spectrum. We analyze the trade‐offs and limitations of this design, and demonstrate its use in a camera, projector and light source. We propose applications such as adaptive color primaries, metamer detection, scene contrast enhancement, photographing fluorescent objects, and high dynamic range photography using spectrum modulation. Ankit Mohan, Ramesh Raskar, Jack Tumblin |
Comput. Graph. Forum | 2 |
| 2008 | Multiflash Stereopsis: Depth-Edge-Preserving Stereo with Small Baseline IlluminationabstractTraditional stereo matching algorithms are limited in their ability to produce accurate results near depth discontinuities, due to partial occlusions and violation of smoothness constraints. In this paper, we use small baseline multi-flash illumination to produce a rich set of feature maps that enable acquisition of discontinuity preserving point correspondences. First, from a single multi-flash camera, we formulate a qualitative depth map using a gradient domain method that encodes object relative distances. Then, in a multiview setup, we exploit shadows created by light sources to compute an occlusion map. Finally, we demonstrate the usefulness of these feature maps by incorporating them into two different dense stereo correspondence algorithms, the first based on local search and the second based on belief propagation. Experimental results show that our enhanced stereo algorithms are able to extract high quality, discontinuity preserving correspondence maps from scenes that are extremely challenging for conventional stereo methods. We also demonstrate that small baseline illumination can be useful to handle specular reflections in stereo imagery. Different from most existing active illumination techniques, our method is simple, inexpensive, compact, and requires no calibration of light sources. Rogério Feris, Ramesh Raskar, Longbin Chen, Kar-Han Tan, Matthew Turk 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Towards passive 6D reflectance field displaysabstractTraditional flat screen displays present 2D images. 3D and 4D displays have been proposed making use of lenslet arrays to shape a fixed outgoing light field for horizontal or bidirectional parallax. In this article, we present different designs of multi-dimensional displays which passively react to the light of the environment behind. The prototypes physically implement a reflectance field and generate different light fields depending on the incident illumination, for example light falling through a window. We discretize the incident light field using an optical system, and modulate it with a 2D pattern, creating a flat display which is view and illumination-dependent. It is free from electronic components. For distant light and a fixed observer position, we demonstrate a passive optical configuration which directly renders a 4D reflectance field in the real-world illumination behind it. We further propose an optical setup that allows for projecting out different angular distributions depending on the incident light direction. Combining multiple of these devices we build a display that renders a 6D experience, where the incident 2D illumination influences the outgoing light field, both in the spatial and in the angular domain. Possible applications of this technology are time-dependent displays driven by sunlight, object virtualization and programmable light benders / ray blockers without moving parts. Martin Fuchs 0001, Ramesh Raskar, Hans-Peter Seidel, Hendrik P. A. Lensch |
ACM Trans. Graph. | 2 |
| 2008 | Shield fields: modeling and capturing 3D occludersabstractWe describe a unified representation of occluders in light transport and photography using shield fields: the 4D attenuation function which acts on any light field incident on an occluder. Our key theoretical result is that shield fields can be used to decouple the effects of occluders and incident illumination. We first describe the properties of shield fields in the frequency-domain and briefly analyze the "forward" problem of efficiently computing cast shadows. Afterwards, we apply the shield field signal-processing framework to make several new observations regarding the "inverse" problem of reconstructing 3D occluders from cast shadows -- extending previous work on shape-from-silhouette and visual hull methods. From this analysis we develop the first single-camera, single-shot approach to capture visual hulls without requiring moving or programmable illumination. We analyze several competing camera designs, ultimately leading to the development of a new large-format, mask-based light field camera that exploits optimal tiled-broadband codes for light-efficient shield field capture. We conclude by presenting a detailed experimental analysis of shield field capture and 3D occluder reconstruction. Douglas Lanman, Ramesh Raskar, Amit K. Agrawal, Gabriel Taubin |
ACM Trans. Graph. | 2 |
| 2008 | Glare aware photography: 4D ray sampling for reducing glare effects of camera lensesabstractGlare arises due to multiple scattering of light inside the camera's body and lens optics and reduces image contrast. While previous approaches have analyzed glare in 2D image space, we show that glare is inherently a 4D ray-space phenomenon. By statistically analyzing the ray-space inside a camera, we can classify and remove glare artifacts. In ray-space, glare behaves as high frequency noise and can be reduced by outlier rejection. While such analysis can be performed by capturing the light field inside the camera, it results in the loss of spatial resolution. Unlike light field cameras, we do not need to reversibly encode the spatial structure of the ray-space, leading to simpler designs. We explore masks for uniform and non-uniform ray sampling and show a practical solution to analyze the 4D statistics without significantly compromising image resolution. Although diffuse scattering of the lens introduces 4D low-frequency glare, we can produce useful solutions in a variety of common scenarios. Our approach handles photography looking into the sun and photos taken without a hood, removes the effect of lens smudges and reduces loss of contrast due to camera body reflections. We show various applications in contrast enhancement and glare manipulation. Ramesh Raskar, Amit K. Agrawal, Cyrus A. Wilson, Ashok Veeraraghavan |
ACM Trans. Graph. | 1 |
| 2007 | Detecting and Segmenting Un-occluded Items by Actively Casting Shadows
Tze Ki Koh, Amit K. Agrawal, Ramesh Raskar, Steve Morgan, Nicholas Miles, Barrie Hayes-Gill |
ACCV (1) | 3 |
| 2007 | Less Is More: Coded Computational Photography
Ramesh Raskar |
ACCV (1) | 1 |
| 2007 | Resolving Objects at Higher Resolution from a Single Motion-blurred ImageabstractMotion blur can degrade the quality of images and is considered a nuisance for computer vision problems. In this paper, we show that motion blur can in-fact be used for increasing the resolution of a moving object. Our approach utilizes the information in a single motion-blurred image without any image priors or training images. As the blur size increases, the resolution of the moving object can be enhanced by a larger factor, albeit with a corresponding increase in reconstruction noise. Traditionally, motion deblurring and super-resolution have been ill-posed problems. Using a coded-exposure camera that preserves high spatial frequencies in the blurred image, we present a linear algorithm for the combined problem of deblurring and resolution enhancement and analyze the invertibility of the resulting linear system. We also show a method to selectively enhance the resolution of a narrow region of high-frequency features, when the resolution of the entire moving object cannot be increased due to small motion blur. Results on real images showing up to four times resolution enhancement are presented. Amit K. Agrawal, Ramesh Raskar |
CVPR | 2 |
| 2007 | Visual Chatter in the Real World
Shree K. Nayar, Gurunandan Krishnan, Michael D. Grossberg, Ramesh Raskar |
ISRR | 4 |
| 2007 | Videoshop: A new framework for spatio-temporal video editing in gradient domain
Ning Xu 0005, Ramesh Raskar, Narendra Ahuja |
Graph. Model. | 3 |
| 2007 | Prakash: lighting aware motion capture using photosensing markers and multiplexed illuminatorsabstractIn this paper, we present a high speed optical motion capture method that can measure three dimensional motion, orientation, and incident illumination at tagged points in a scene. We use tracking tags that work in natural lighting conditions and can be imperceptibly embedded in attire or other objects. Our system supports an unlimited number of tags in a scene, with each tag uniquely identified to eliminate marker reacquisition issues. Our tags also provide incident illumination data which can be used to match scene lighting when inserting synthetic elements. The technique is therefore ideal for on-set motion capture or real-time broadcasting of virtual sets. Unlike previous methods that employ high speed cameras or scanning lasers, we capture the scene appearance using the simplest possible optical devices - a light-emitting diode (LED) with a passive binary mask used as the transmitter and a photosensor used as the receiver. We strategically place a set of optical transmitters to spatio-temporally encode the volume of interest. Photosensors attached to scene points demultiplex the coded optical signals from multiple transmitters, allowing us to compute not only receiver location and orientation but also their incident illumination and the reflectance of the surfaces to which the photosensors are attached. We use our untethered tag system, called Prakash, to demonstrate methods of adding special effects to captured videos that cannot be accomplished using pure vision techniques that rely on camera images. Ramesh Raskar, Hideaki Nii, Bert de Decker, Yuki Hashimoto, Jay Summet, Dylan Moore 0002, Jonathan Westhues, Paul H. Dietz, John Barnwell, Shree K. Nayar, Masahiko Inami, Philippe Bekaert, Michael Noland, Vlad Branzoi, Erich Bruns |
ACM Trans. Graph. | 1 |
| 2007 | Dappled photography: mask enhanced cameras for heterodyned light fields and coded aperture refocusingabstractWe describe a theoretical framework for reversibly modulating 4D light fields using an attenuating mask in the optical path of a lens based camera. Based on this framework, we present a novel design to reconstruct the 4D light field from a 2D camera image without any additional refractive elements as required by previous light field cameras. The patterned mask attenuates light rays inside the camera instead of bending them, and the attenuation recoverably encodes the rays on the 2D sensor. Our mask-equipped camera focuses just as a traditional camera to capture conventional 2D photos at full sensor resolution, but the raw pixel values also hold a modulated 4D light field. The light field can be recovered by rearranging the tiles of the 2D Fourier transform of sensor values into 4D planes, and computing the inverse Fourier transform. In addition, one can also recover the full resolution image information for the in-focus parts of the scene. We also show how a broadband mask placed at the lens enables us to compute refocused images at full sensor resolution for layered Lambertian scenes. This partial encoding of 4D ray-space data enables editing of image contents by depth, yet does not require computational recovery of the complete 4D light field. Ashok Veeraraghavan, Ramesh Raskar, Amit K. Agrawal, Ankit Mohan, Jack Tumblin |
ACM Trans. Graph. | 2 |
| 2006 | A Handheld Projector Supported by Computer Vision
Akash Kushal, Jeroen van Baar, Ramesh Raskar, Paul A. Beardsley |
ACCV (2) | 3 |
| 2006 | Edge Suppression by Gradient Field Transformation Using Cross-Projection TensorsabstractWe propose a new technique for edge-suppressing operations on images. We introduce cross projection tensors to achieve affine transformations of gradient fields. We use these tensors, for example, to remove edges in one image based on the edge-information in a second image. Traditionally, edge suppression is achieved by setting image gradients to zero based on thresholds. A common application is in the Retinex problem, where the illumination map is recovered by suppressing the reflectance edges, assuming it is slowly varying. We present a class of problems where edge-suppression can be a useful tool. These problems involve analyzing images of the same scene under variable illumination. Instead of resetting gradients, the key idea in our approach is to derive local tensors using one image and to transform the gradient field of another image using them. Reconstructed image from the modified gradient field shows suppressed edges or textures at the corresponding locations. All operations are local and our approach does not require any global analysis. We demonstrate the algorithm in the context of several applications such as (a) recovering the foreground layer under varying illumination, (b) estimating intrinsic images in non-Lambertian scenes, (c) removing shadows from color images and obtaining the illumination map, and (d) removing glass reflections. Amit K. Agrawal, Ramesh Raskar, Rama Chellappa |
CVPR (2) | 2 |
| 2006 | What Is the Range of Surface Reconstructions from a Gradient Field?
Amit K. Agrawal, Ramesh Raskar, Rama Chellappa |
ECCV (1) | 2 |
| 2006 | Keynote speeches - The poor man's palace: special effects in the real world / Do we have six brains?abstractProvides an abstract for each of the keynote presentations and may include a brief professional biography of each Ramesh Raskar, Tom A. Furness |
ISMAR | 1 |
| 2006 | Fast separation of direct and global components of a scene using high frequency illuminationabstractWe present fast methods for separating the direct and global illumination components of a scene measured by a camera and illuminated by a light source. In theory, the separation can be done with just two images taken with a high frequency binary illumination pattern and its complement. In practice, a larger number of images are used to overcome the optical and resolution limitations of the camera and the source. The approach does not require the material properties of objects and media in the scene to be known. However, we require that the illumination frequency is high enough to adequately sample the global components received by scene points. We present separation results for scenes that include complex interreflections, subsurface scattering and volumetric scattering. Several variants of the separation approach are also described. When a sinusoidal illumination pattern is used with different phase shifts, the separation can be done using just three images. When the computed images are of lower resolution than the source and the camera, smoothness constraints are used to perform the separation using a single image. Finally, in the case of a static scene that is lit by a simple point source, such as the sun, a moving occluder and a video camera can be used to do the separation. We also show several simple examples of how novel images of a scene can be computed from the separation results. Shree K. Nayar, Gurunandan Krishnan, Michael D. Grossberg, Ramesh Raskar |
ACM Trans. Graph. | 4 |
| 2006 | Coded exposure photography: motion deblurring using fluttered shutterabstractIn a conventional single-exposure photograph, moving objects or moving cameras cause motion blur. The exposure time defines a temporal box filter that smears the moving object across the image by convolution. This box filter destroys important high-frequency spatial details so that deblurring via deconvolution becomes an ill-posed problem.Rather than leaving the shutter open for the entire exposure duration, we "flutter" the camera's shutter open and closed during the chosen exposure time with a binary pseudo-random sequence. The flutter changes the box filter to a broad-band filter that preserves high-frequency spatial details in the blurred image and the corresponding deconvolution becomes a well-posed problem. We demonstrate that manually-specified point spread functions are sufficient for several challenging cases of motion-blur removal including extremely large motions, textured backgrounds and partial occluders. Ramesh Raskar, Amit K. Agrawal, Jack Tumblin |
ACM Trans. Graph. | 1 |
| 2005 | Why I Want a Gradient CameraabstractWe propose a camera that measures static gradients instead of static intensities. Quantizing sensed intensity differences between adjacent pixel values permits an ordinary A/D converter to measure detailed high contrast (HDR) scenes. We measure alternating 'cliques' of sensors (small groups) that locally determine their own best exposure, and reconstruct the image using a Poisson solver. This intrinsically differential design suppresses common-mode noise, hides and smoothes quantization, and can correct for its own saturated sensors. Simulations demonstrate these capabilities in side-by-side comparisons. Jack Tumblin, Amit K. Agrawal, Ramesh Raskar |
CVPR (1) | 3 |
| 2005 | Videoshop: A New Framework for Spatio-Temporal Video Editing in Gradient DomainabstractOur goal is to develop tools that go beyond frame-constrained manipulation such as resizing, color correction, and simple transitions, and provide object-level operations within frames. Some of our targeted video editing tasks includes transferring a motion picture to a new still picture, importing a moving object into a new background, and compositing two video sequences. The challenges behind this kind of complex video editing tasks lie in two constraints: 1) Spatial consistency: imported objects should blend with the background seamlessly. Hence pixel replacement, which creates noticeable seams, is problematic. 2) Temporal coherency: successive frames should display smooth transitions. Hence frame-by-frame editing, which results in visual flicker, is inappropriate. Our work is aimed at providing an easy-to-use video editing tool that maximally satisfies the spatial and temporal constraints mentioned above and requires minimum user interaction. We propose a new framework for video editing in gradient domain. The spatio-temporal gradient fields of target videos are modified and/or mixed to generate a new gradient field which is usually not integrable. We propose a 3D video integration algorithm, which uses the variational method, to find the potential function whose gradient field is closest to the mixed gradient field in the sense of least squares. The video is reconstructed by solving a 3D Poisson equation. We derive an extension of current 2D gradient technique to 3D space, yielding in a novel video editing framework, which is very different from all current video editing software. Ning Xu 0005, Ramesh Raskar, Narendra Ahuja |
CVPR (2) | 3 |
| 2005 | An Algebraic Approach to Surface Reconstruction from Gradient FieldsabstractSeveral important problems in computer vision such as shape from shading (SFS) and photometric stereo (PS) require reconstructing a surface from an estimated gradient field, which is usually non-integrable, i.e. have non-zero curl. We propose a purely algebraic approach to enforce integrability in discrete domain. We first show that enforcing integrability can be formulated as solving a single linear system Ax =b over the image. In general, this system is under-determined. We show conditions under which the system can be solved and a method to get to those conditions based on graph theory. The proposed approach is non-iterative, has the important property of local error confinement and can be applied to several problems. Results on SFS and PS demonstrate the applicability of our method. Amit K. Agrawal, Rama Chellappa, Ramesh Raskar |
ICCV | 3 |
| 2005 | Discontinuity Preserving Stereo with Small Baseline Multi-Flash IlluminationabstractCurrently, sharp discontinuities in depth and partial occlusions in multiview imaging systems pose serious challenges for many dense correspondence algorithms. However, it is important for 3D reconstruction methods to preserve depth edges as they correspond to important shape features like silhouettes which are critical for understanding the structure of a scene. In this paper, we show how active illumination algorithms can produce a rich set of feature maps that are useful in dense 3D reconstruction. We start by showing a method to compute a qualitative depth map from a single camera, which encodes object relative distances and can be used as a prior for stereo. In a multiview setup, we show that along with depth edges, binocular half-occluded pixels can also be explicitly and reliably labeled. To demonstrate the usefulness of these feature maps, we show how they can be used in two different algorithms for dense stereo correspondence. Our experimental results show that our enhanced stereo algorithms are able to extract high quality, discontinuity preserving correspondence maps from scenes that are extremely challenging for conventional stereo methods. Rogério Feris, Ramesh Raskar, Longbin Chen, Kar-Han Tan, Matthew Turk 0001 |
ICCV | 2 |
| 2005 | Automatic image retargetingabstractWe present a non-photorealistic algorithm for retargeting large images to small size displays, particularly on mobile devices. This method adapts large images so that important objects in the image are still recognizable when displayed at a lower target resolution. Existing image manipulation techniques such as cropping works well for images containing a single important object, and down-sampling works well for images containing low frequency information. However, when these techniques are automatically applied to images with multiple objects, the image quality degrades and important information may be lost. Our algorithm addresses the case of multiple important objects in an image. The retargeting algorithm segments an image into regions, identifies important regions, removes them, fills the resulting gaps, resizes the remaining image, and re-inserts the important regions. Our approach lies in constructing a topologically constrained epitome of an image based on a visual attention model that is both comprehensible and size varying, making the method suitable for display-critical applications. Vidya Setlur, Saeko Takagi, Ramesh Raskar, Michael Gleicher, Bruce Gooch |
MUM | 3 |
| 2005 | Zoom-and-pick: facilitating visual zooming and precision pointing with interactive handheld projectorsabstractDesigning interfaces for interactive handheld projectors is an exiting new area of research that is currently limited by two problems: hand jitter resulting in poor input control, and possible reduction of image resolution due to the needs of image stabilization and warping algorithms. We present the design and evaluation of a new interaction technique, called zoom-and-pick, that addresses both problems by allowing the user to fluidly zoom in on areas of interest and make accurate target selections. Subtle design features of zoom-and-pick enable pixel-accurate pointing, which is not possible in most freehand interaction techniques. Our evaluation results indicate that zoom-and-pick is significantly more accurate than the standard pointing technique described in our previous work. Clifton Forlines, Ravin Balakrishnan, Paul A. Beardsley, Jeroen van Baar, Ramesh Raskar |
UIST | 5 |
| 2005 | Gradient domain context enhancement for fixed camerasabstractWe propose a class of enhancement techniques suitable for scenes captured by fixed cameras. The basic idea is to increase the information density in a set of low quality images by extracting the context from a higher-quality image captured under different illuminations from the same viewpoint. For example, a night-time surveillance video can be enriched with information available in daytime images. We also propose a new image fusion approach to combine images with sufficiently different appearance into a seamless rendering. Our method ensures the fidelity of important features and robustly incorporates background contexts, while avoiding traditional problems such as aliasing, ghosting and haloing. We show results on indoor as well as outdoor scenes. Adrian Ilie, Ramesh Raskar, Jingyi Yu 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2005 | Removing photography artifacts using gradient projection and flash-exposure samplingabstractFlash images are known to suffer from several problems: saturation of nearby objects, poor illumination of distant objects, reflections of objects strongly lit by the flash and strong highlights due to the reflection of flash itself by glossy surfaces. We propose to use a flash and no-flash (ambient) image pair to produce better flash images. We present a novel gradient projection scheme based on a gradient coherence model that allows removal of reflections and highlights from flash images. We also present a brightness-ratio based algorithm that allows us to compensate for the falloff in the flash image brightness due to depth. In several practical scenarios, the quality of flash/no-flash images may be limited in terms of dynamic range. In such cases, we advocate using several images taken under different flash intensities and exposures. We analyze the flash intensity-exposure space and propose a method for adaptively sampling this space so as to minimize the number of captured images for any given scene. We present several experimental results that demonstrate the ability of our algorithms to produce improved flash images. Amit K. Agrawal, Ramesh Raskar, Shree K. Nayar, Yuanzhen Li |
ACM Trans. Graph. | 2 |
| 2004 | Multi-projectors and implicit interaction in persuasive public displaysabstractRecent advances in computer video projection open up new possibilities for real-time interactive, persuasive displays. Now a display can continuously adapt to a viewer so as to maximize its effectiveness. However, by the very nature of persuasion, these displays must be both immersive and subtle. We have been working on technologies that support this application including multi-projector and implicit interaction techniques. These technologies have been used to create a series of interactive persuasive displays that are described. Paul H. Dietz, Ramesh Raskar, Shane Booth, Jeroen van Baar, Kent Wittenburg, Brian Knep |
AVI | 2 |
| 2004 | Shape-Enhanced Surgical Visualizations and Medical Illustrations with Multi-flash Imaging
Kar-Han Tan, James Kobler, Paul H. Dietz, Ramesh Raskar, Rogério Feris |
MICCAI (2) | 4 |
| 2004 | Automatic projector calibration with embedded light sensorsabstractProjection technology typically places several constraints on the geometric relationship between the projector and the projection surface to obtain an undistorted, properly sized image. In this paper we describe a simple, robust, fast, and low-cost method for automatic projector calibration that eliminates many of these constraints. We embed light sensors in the target surface, project Gray-coded binary patterns to discover the sensor locations, and then prewarp the image to accurately fit the physical features of the projection surface. This technique can be expanded to automatically stitch multiple projectors, calibrate onto non-planar surfaces for object decoration, and provide a method for simple geometry acquisition. Johnny C. Lee, Paul H. Dietz, Dan Maynes-Aminzade, Ramesh Raskar, Scott E. Hudson |
UIST | 4 |
| 2004 | Quadric Transfer for Immersive Curved Screen DisplaysabstractAbstract Curved screens are increasingly being used for high‐resolution immersive visualization environments. We describe a new technique to display seamless images using overlapping projectors on curved quadric surfaces such as spherical or cylindrical shape. We exploit a quadric image transfer function and show how it can be used to achieve sub‐pixel registration while interactively displaying two or three‐dimensional datasets for a head‐tracked user. Current techniques for automatically registered seamless displays have focused mainly on planar displays. On the other hand, techniques for curved screens currently involve cumbersome manual alignment to make the installation conform to the intended design. We show a seamless real‐time display system and discuss our methods for smooth intensity blending and efficient rendering. Categories and Subject Descriptors (according to ACM CCS): I.3.7 [Computer Graphics]: Three‐Dimensional Graphics and Realism‐ Virtual reality Ramesh Raskar, Jeroen van Baar, Thomas Willwacher, Srinivas Rao |
Comput. Graph. Forum | 1 |
| 2004 | RFIG lamps: interacting with a self-describing world via photosensing wireless tags and projectorsabstractThis paper describes how to instrument the physical world so that objects become self-describing, communicating their identity, geometry, and other information such as history or user annotation. The enabling technology is a wireless tag which acts as a radio frequency identity and geometry (RFIG) transponder. We show how addition of a photo-sensor to a wireless tag significantly extends its functionality to allow geometric operations - such as finding the 3D position of a tag, or detecting change in the shape of a tagged object. Tag data is presented to the user by direct projection using a handheld locale-aware mobile projector. We introduce a novel technique that we call interactive projection to allow a user to interact with projected information e.g. to navigate or update the projected information.The ideas are demonstrated using objects with active radio frequency (RF) tags. But the work was motivated by the advent of unpowered passive-RFID, a technology that promises to have significant impact in real-world applications. We discuss how our current prototypes could evolve to passive-RFID in the future. Ramesh Raskar, Paul A. Beardsley, Jeroen van Baar, Paul H. Dietz, Johnny C. Lee, Darren Leigh, Thomas Willwacher |
ACM Trans. Graph. | 1 |
| 2004 | Non-photorealistic camera: depth edge detection and stylized rendering using multi-flash imagingabstractWe present a non-photorealistic rendering approach to capture and convey shape features of real-world scenes. We use a camera with multiple flashes that are strategically positioned to cast shadows along depth discontinuities in the scene. The projective-geometric relationship of the camera-flash setup is then exploited to detect depth discontinuities and distinguish them from intensity edges due to material discontinuities.We introduce depiction methods that utilize the detected edge features to generate stylized static and animated images. We can highlight the detected features, suppress unnecessary details or combine features from multiple images. The resulting images more clearly convey the 3D structure of the imaged scenes.We take a very different approach to capturing geometric features of a scene than traditional approaches that require reconstructing a 3D model. This results in a method that is both surprisingly simple and computationally efficient. The entire hardware/software setup can conceivably be packaged into a self-contained device no larger than existing digital cameras. Ramesh Raskar, Kar-Han Tan, Rogério Feris, Jingyi Yu 0001, Matthew Turk 0001 |
ACM Trans. Graph. | 1 |
| 2003 | iLamps: geometrically aware and self-configuring projectorsabstractProjectors are currently undergoing a transformation as they evolve from static output devices to portable, environment-aware, communicating systems. An enhanced projector can determine and respond to the geometry of the display surface, and can be used in an ad-hoc cluster to create a self-configuring display. Information display is such a prevailing part of everyday life that new and more flexible ways to present data are likely to have significant impact. This paper examines geometrical issues for enhanced projectors, relating to customized projection for different shapes of display surface, object augmentation, and co-operation between multiple units.We introduce a new technique for adaptive projection on nonplanar surfaces using conformal texture mapping. We describe object augmentation with a hand-held projector, including interaction techniques. We describe the concept of a display created by an ad-hoc cluster of heterogeneous enhanced projectors, with a new global alignment scheme, and new parametric image transfer methods for quadric surfaces, to make a seamless projection. The work is illustrated by several prototypes and applications. Ramesh Raskar, Jeroen van Baar, Paul A. Beardsley, Thomas Willwacher, Srinivas Rao, Clifton Forlines |
ACM Trans. Graph. | 1 |
| 2002 | Blending Multiple ViewsabstractCurrent blending methods in image-based rendering use local information such as "deviations from the closest views" to find blending weights. They include approaches such as the view-dependent texture mapping and blending fields used in unstructured lumigraph rendering. However, in the presence of depth discontinuities, these techniques do not provide smooth transitions in the target image if the intensities of corresponding pixels in the source images are significantly different (e.g. due to specular highlights). In this paper, we present an image blending technique that allows the use of global visibility and occlusion constraints. Each blending weight now has a global component and a local component, which are due to the view-independent and view-dependent contributions of the source images, respectively. Being view-independent, the global components can be computed in a preprocessing stage. Traditional graphics hardware is exploited to accelerate the computation of the global blending weights. Ramesh Raskar, Kok-Lim Low |
PG | 1 |
| 2002 | Free-form sketching with variational implicit surfacesabstractWith the advent of sketch-based methods for shape construction, there's a new degree of power available in the rapid creation of approximate shapes. Sketch [Zeleznik, 1996] showed how a gesture-based modeler could be used to simplify conventional CSG-like shape creation. Teddy [Igarashi, 1999] extended this to more free-form models, getting much of its power from its ``inflation'' operation (which converted a simple closed curve in the plane into a 3D shape whose silhouette, from the current point of view, was that curve on the view plane) and from an elegant collection of gestures for attaching additional parts to a shape, cutting a shape, and deforming it. But despite the powerful collection of tools in Teddy, the underlying polygonal representation of shapes intrudes on the results in many places. In this paper, we discuss our preliminary efforts at using variational implicit surfaces [Turk, 2000] as a representation in a free-form modeler. We also discuss the implementation of several operations within this context, and a collection of user-interaction elements that work well together to make modeling interesting hierarchies simple. These include ``stroke inflation'' via implicit functions, blob-merging, automatic hierarchy construction, and local surface modification via silhouette oversketching. We demonstrate our results by creating several models. Categories and Subject Descriptors (according to ACM CCS): I.3.5 [Computer Graphics]: Modeling packages I.3.6 [Computer Graphics]: Interaction techniques Olga A. Karpenko, John F. Hughes, Ramesh Raskar |
Comput. Graph. Forum | 3 |
| 2001 | A Self-Correcting ProjectorabstractWe describe a calibration and rendering technique for a projector that can render rectangular images under keystoned position. The projector utilizes a rigidly attached camera to form a stereo pair. We describe a very easy to use technique for calibration of the projector-camera pair using only black planar surfaces. We present an efficient rendering method to pre-warp images so that they appear correctly on the screen, and show experimental results. Ramesh Raskar, Paul A. Beardsley |
CVPR (2) | 1 |
| 2000 | Image-based visual hullsabstractIn this paper, we describe an efficient image-based approach to computing and shading visual hulls from silhouette image data. Our algorithm takes advantage of epipolar geometry and incremental computation to achieve a constant rendering cost per rendered pixel. It does not suffer from the computation complexity, limited resolution, or quantization artifacts of previous volumetric approaches. We demonstrate the use of this algorithm in a real-time virtualized reality application running off a small number of video streams. Keywords: Computer Vision, Image-Based Rendering, Constructive Solid Geometry, Misc. Rendering Algorithms. 1 Introduction Visualizing and navigating within virtual environments composed of both real and synthetic objects has been a long-standing goal of computer graphics. The term "Virtualized Reality^TM", as popularized by Kanade [23], describes a setting where a real-world scene is "captured" by a collection of cameras and then viewed through a virtual camera, a... Wojciech Matusik, Chris Buehler, Ramesh Raskar, Steven J. Gortler, Leonard McMillan |
SIGGRAPH | 3 |
| 2000 | Immersive Planar Display using Roughly Aligned ProjectorsabstractWhen a projector is oblique with respect to a planar display surface, it creates keystoning and the projected image is distorted. We present a rendering technique to display perspectively correct images for a moving user. This allows using roughly aligned projectors and eliminates the need for frequent electro-mechanical adjustments. The rendering process has no additional cost and can be implemented with traditional graphics hardware. We compute the collineation induced due to the display plane during preprocessing. The main idea of the paper is to use this collineation to render and warp the images of 3D scenes in a single pass via approximation of the depth buffer. We also describe how this method can be extended to display systems with multiple overlapping projectors. This technique can be easily used in CAVE, Immersive Workbenches and PowerWalls. Ramesh Raskar |
VR | 1 |
| 1999 | Image precision silhouette edgesabstractArticle Image precision silhouette edges Share on Authors: Ramesh Raskar University of North Carolina at Chapel Hill University of North Carolina at Chapel HillView Profile , Michael Cohen Microsoft Research Microsoft ResearchView Profile Authors Info & Claims I3D '99: Proceedings of the 1999 symposium on Interactive 3D graphicsApril 1999 Pages 135–140https://doi.org/10.1145/300523.300539Online:26 April 1999Publication History 107citation1,038DownloadsMetricsTotal Citations107Total Downloads1,038Last 12 Months30Last 6 weeks9 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Ramesh Raskar, Michael F. Cohen |
SI3D | 1 |
| 1999 | Multi-Projector Displays Using Camera-Based RegistrationabstractConventional projector-based display systems are typically designed around precise and regular configurations of projectors and display surfaces. While this results in rendering simplicity and speed, it also means painstaking construction and ongoing maintenance. In previously published work, we introduced a vision of projector-based displays constructed from a collection of casually-arranged projectors and display surfaces. In this paper, we present flexible yet practical methods for realizing this vision, enabling low-cost mega-pixel display systems with large physical dimensions, higher resolution, or both. The techniques afford new opportunities to build personal 3D visualization systems in offices, conference rooms, theaters, or even your living room. As a demonstration of the simplicity and effectiveness of the methods that we continue to perfect, we show in the included video that a 10-year old child can construct and calibrate a two-camera, two-projector, head-tracked display system, all in about 15 minutes. Ramesh Raskar, Michael S. Brown, Ruigang Yang, Wei-Chao Chen, Greg Welch, Herman Towles, W. Brent Seales, Henry Fuchs |
IEEE Visualization | 1 |
| 1998 | Augmented Reality Visualization for Laparoscopic Surgery
Henry Fuchs, Mark A. Livingston, Ramesh Raskar, D'nardo Colucci, Kurtis Keller, Andrei State, Jessica R. Crawford, Paul Rademacher, Samuel H. Drake, Anthony A. Meyer |
MICCAI | 3 |
| 1998 | The Office of the Future: A Unified Approach to Image-based Modeling and Spatially Immersive DisplaysabstractWe introduce ideas, proposed technologies, and initial results for an office of the future that is based on a unified application of computer vision and computer graphics in a system that combines and builds upon the notions of the CAVE™, tiled display systems, and image-based modeling .The basic idea is to use real-time computer vision techniques to dynamically extract per-pixel depth and reflectance information for the visible surfaces in the office including walls, furniture, objects, and people, and then to either project images on the surfaces, render images of the surfaces , or interpret changes in the surfaces.In the first case, one could designate every-day (potentially irregular) real surfaces in the office to be used as spatially immersive display surfaces, and then project high-resolution graphics and text onto those surfaces.In the second case, one could transmit the dynamic image-based models over a network for display at a remote site.Finally, one could interpret dynamic changes in the surfaces for the purposes of tracking, interaction, or augmented reality applications.To accomplish the simultaneous capture and display we envision an office of the future where the ceiling lights are replaced by computer controlled cameras and "smart" projectors that are used to capture dynamic image-based models with imperceptible structured light techniques, and to display high-resolution images on designated display surfaces.By doing both simultaneously on the designated display surfaces, one can dynamically adjust or autocalibrate for geometric, intensity, and resolution variations resulting from irregular or changing display surfaces, or overlapped projector images.Our current approach to dynamic image-based modeling is to use an optimized structured light scheme that can capture per-pixel depth and reflectance at interactive rates.Our system implementation is not yet imperceptible, but we can demonstrate the approach in the laboratory.Our approach to rendering on the designated (potentially irregular) display surfaces is to employ a two-pass projective texture scheme to generate images that when projected onto the surfaces appear correct to a moving headtracked observer.We present here an initial implementation of the overall vision, in an office-like setting, and preliminary demonstrations of our dynamic modeling and display techniques. Ramesh Raskar, Greg Welch, Matthew D. Cutts, Adam T. Lake, Lev Stesin, Henry Fuchs |
SIGGRAPH | 1 |