Akshat Dave

dblp:126/0916 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0003-0560-632XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Blurred LiDAR for Sharper 3D: Robust Handheld 3D Scanning with Diffuse LiDAR and RGB
abstract
3D surface reconstruction is essential across applications of virtual reality, robotics, and mobile scanning. However, RGB-based reconstruction often fails in low-texture, low-light, and low-albedo scenes. Handheld LiDARs, now common on mobile devices, aim to address these challenges by capturing depth information from time-of-flight measurements of a coarse grid of projected dots. Yet, these sparse LiDARs struggle with scene coverage on limited input views, leaving large gaps in depth information. In this work, we propose using an alternative class of "blurred" LiDAR that emits a diffuse flash, greatly improving scene coverage but introducing spatial ambiguity from mixed time-of-flight measurements across a wide field of view. To handle these ambiguities, we propose leveraging the complementary strengths of diffuse LiDAR with RGB. We introduce a Gaussian surfel-based rendering framework with a scene-adaptive loss function that dynamically balances RGB and diffuse LiDAR signals. We demonstrate that, surprisingly, diffuse LiDAR can outperform traditional sparse LiDAR, enabling robust 3D scanning with accurate color and geometry estimation in challenging environments.
Nikhil Behari, Aaron Young, Siddharth Somasundaram, Tzofi Klinghoffer, Akshat Dave, Ramesh Raskar
CVPR5
2025 Enhancing Autonomous Navigation by Imaging Hidden Objects Using Single-Photon LiDAR
abstract
Robust autonomous navigation in environments with limited visibility remains a critical challenge in robotics. We present a novel approach that leverages Non-Line-of-Sight (NLOS) sensing using single-photon LiDAR to improve visibility and enhance autonomous navigation. Our method enables mobile robots to “see around corners” by utilizing multi-bounce light information, effectively expanding their perceptual range without additional infrastructure. We propose a three-module pipeline: (1) Sensing, which captures multi-bounce histograms using SPAD-based LiDAR; (2) Perception, which estimates occupancy maps of hidden regions from these histograms using a convolutional neural network; and (3) Control, which allows a robot to follow safe paths based on the estimated occupancy. We evaluate our approach through simulations and real-world experiments on a mobile robot navigating an L-shaped corridor with hidden obstacles. Our work represents the first experimental demonstration of NLOS imaging for autonomous navigation, paving the way for safer and more efficient robotic systems operating in complex environments. We also contribute a novel dynamics-integrated transient rendering framework for simulating NLOS scenarios, facilitating future research in this domain.
Aaron Young, Nevindu Batagoda, Harry Zhang, Akshat Dave, Adithya Kumar Pediredla, Dan Negrut, Ramesh Raskar
ICRA4
2025 Shoot-Bounce-3D: Single-Shot Occlusion-Aware 3D from Lidar by Decomposing Two-Bounce Light
abstract
3D scene reconstruction from a single measurement is challenging, especially in the presence of occluded regions and specular materials, such as mirrors. We address these challenges by leveraging single-photon lidars. These lidars estimate depth from light that is emitted into the scene and reflected directly back to the sensor. However, they can also measure light that bounces multiple times in the scene before reaching the sensor. This multi-bounce light contains additional information that can be used to recover dense depth, occluded geometry, and material properties. Prior work with single-photon lidar, however, has only demonstrated these use cases when a laser sequentially illuminates one scene point at a time. We instead focus on the more practical – and challenging – scenario of illuminating multiple scene points simultaneously. The complexity of light transport due to the combined effects of multiplexed illumination, two-bounce light, shadows, and specular reflections is challenging to invert analytically. Instead, we propose a data-driven method to invert light transport in single-photon lidar. To enable this approach, we create the first large-scale simulated dataset of ~100k lidar transients for indoor scenes. We use this dataset to learn a prior on complex light transport, enabling measured two-bounce light to be decomposed into the constituent contributions from each laser spot. Finally, we experimentally demonstrate how this decomposed light can be used to infer 3D geometry in scenes with occlusions and mirrors from a single measurement. Our code and dataset are released on our project webpage.
Tzofi Klinghoffer, Siddharth Somasundaram, Xiaoyu Xiang, Yuchen Fan 0001, Christian Richardt, Akshat Dave, Ramesh Raskar
SIGGRAPH Asia6
2025 Event Cameras Meet SPADs for High-Speed, Low-Bandwidth Imaging
abstract
Traditional cameras face a trade-off between low-light performance and high-speed imaging: longer exposure times to capture sufficient light results in motion blur, whereas shorter exposures result in Poisson-corrupted noisy images. While burst photography techniques help mitigate this tradeoff, conventional cameras are fundamentally limited in their sensor noise characteristics. Event cameras and single-photon avalanche diode (SPAD) sensors have emerged as promising alternatives to conventional cameras due to their desirable properties. SPADs are capable of single-photon sensitivity with microsecond temporal resolution, and event cameras can measure brightness changes up to 1 MHz with low bandwidth requirements. We show that these properties are complementary, and can help achieve low-light, high-speed image reconstruction with low bandwidth requirements. We introduce a sensor fusion framework to combine SPADs with event cameras to improve the reconstruction of high-speed, low-light scenes while reducing the high bandwidth cost associated with using every SPAD frame. Our evaluation, on both synthetic and real sensor data, demonstrates significant enhancements ($> 5$>5 dB PSNR) in reconstructing low-light scenes at high temporal resolution (100 kHz) compared to conventional cameras. Event-SPAD fusion shows great promise for real-world applications, such as robotics or medical imaging.
Manasi Muglikar, Siddharth Somasundaram, Akshat Dave, Edoardo Charbon, Ramesh Raskar, Davide Scaramuzza 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 NeST: Neural Stress Tensor Tomography by leveraging 3D Photoelasticity
abstract
Photoelasticity enables full-field stress analysis in transparent objects through stress-induced birefringence. Existing techniques are limited to two-dimensional (2D) slices and require destructively slicing the object. Recovering the internal three-dimensional (3D) stress distribution of the entire object is challenging, as it involves solving a tensor tomography problem and handling phase wrapping ambiguities. We introduce NeST, an analysis-by-synthesis approach for reconstructing 3D stress tensor fields as neural implicit representations from polarization measurements. Our key insight is to jointly handle phase unwrapping and tensor tomography using a differentiable forward model based on Jones calculus. Our non-linear model faithfully matches real captures, unlike prior linear approximations. We develop an experimental multi-axis polariscope setup to capture 3D photoelasticity and experimentally demonstrate that NeST reconstructs the internal stress distribution for objects with varying shape and force conditions. Additionally, we showcase novel applications in stress analysis, such as visualizing photoelastic fringes by virtually slicing the object and viewing photoelastic fringes from unseen viewpoints. NeST paves the way for scalable non-destructive 3D photoelastic analysis.
Akshat Dave, Aaron Young, Ramesh Raskar, Wolfgang Heidrich, Ashok Veeraraghavan
ACM Trans. Graph.1
2024 DecentNeRFs: Decentralized Neural Radiance Fields from Crowdsourced Images
Zaid Tasneem, Akshat Dave, Abhishek Singh 0005, Kushagra Tiwary, Praneeth Vepakomma, Ashok Veeraraghavan, Ramesh Raskar
ECCV (59)2
2024 Handheld Mapping of Specular Surfaces Using Consumer-Grade Flash LiDAR
abstract
We propose an approach to leverage multi-bounce returns of a flash LiDAR on portable smartphones for 3D specular surface reconstruction. Traditional LiDAR systems assume that all returns are one-bounce returns, which can lead to an overestimation of the true mirror surface and cause it to appear as if there is a hole. However, in reality, returns from mirror surfaces follow multi-bounce paths. We operate with a consumer-grade, coarse multi-beam flash LiDAR, enabling real-time mapping on an affordable and portable smartphone. To address the challenges posed by the coarse setup, where the transmitter and receiver are co-located, we propose solving the association problem using the ‘reciprocal pair’ algorithm. This algorithm can distinguish between different types of bounces from multi-bounce returns. We have demonstrated detection over multiple consecutive frames for dense mirror mapping. In addition to 3D reconstruction, we show that multi-bounce returns enhance performance in applications such as segmentation and novel view synthesis. Our method can be integrated with state-of-the-art learned-based models, enhancing their robustness in discerning ambiguous scenarios. Importantly, our approach can map various specular surfaces like mirrors and glasses without assuming specific shapes, and it can operate on non-perpendicular specular-diffuse surface pairs.
Tsung-Han Lin, Connor Henley, Siddharth Somasundaram, Akshat Dave, Moshe Laifenfeld, Ramesh Raskar
ICCP4
2023 Role of Transients in Two-Bounce Non-Line-of-Sight Imaging
abstract
The goal of non-line-of-sight (NLOS) imaging is to image objects occluded from the camera's field of view using multiply scattered light. Recent works have demonstrated the feasibility of two-bounce (2B) NLOS imaging by scanning a laser and measuring cast shadows of occluded objects in scenes with two relay surfaces. In this work, we study the role of time-of-flight (ToF) measurements, i.e. transients, in 2B-NLOS under multiplexed illumination. Specifically, we study how ToF information can reduce the number of measurements and spatial resolution needed for shape reconstruction. We present our findings with respect to tradeoffs in (1) temporal resolution, (2) spatial resolution, and (3) number of image captures by studying SNR and recoverability as functions of system parameters. This leads to a formal definition of the mathematical constraints for 2B lidar. We believe that our work lays an analytical ground- workfor design of future NLOS imaging systems, especially as ToF sensors become increasingly ubiquitous.
Siddharth Somasundaram, Akshat Dave, Connor Henley, Ashok Veeraraghavan, Ramesh Raskar
CVPR2
2023 ORCa: Glossy Objects as Radiance-Field Cameras
abstract
Reflections on glossy objects contain valuable and hidden information about the surrounding environment. By converting these objects into cameras, we can unlock exciting applications, including imaging beyond the camera's field-of-view and from seemingly impossible vantage points, e.g. from reflections on the human eye. However, this task is challenging because reflections depend jointly on object geometry, material properties, the 3D environment, and the observer's viewing direction. Our approach converts glossy objects with unknown geometry into radiance-field cameras to image the world from the object's perspective. Our key insight is to convert the object surface into a virtual sensor that captures cast reflections as a 2D projection of the 5D environment radiance field visible to and surrounding the object. We show that recovering the environment radiance fields enables depth and radiance estimation from the object to its surroundings in addition to beyond field-of-view novel-view synthesis, i.e. rendering of novel views that are only directly visible to the glossy object present in the scene, but not the observer. Moreover, using the radiance field we can image around occluders caused by close-by objects in the scene. Our method is trained end-to-end on multi-view images of the object and jointly estimates object geometry, diffuse radiance, and the 5D environment radiance field. For more information, visit our website.
Kushagra Tiwary, Akshat Dave, Nikhil Behari, Tzofi Klinghoffer, Ashok Veeraraghavan, Ramesh Raskar
CVPR2
2022 PANDORA: Polarization-Aided Neural Decomposition of Radiance
Akshat Dave, Yongyi Zhao, Ashok Veeraraghavan
ECCV (7)1
2022 First Arrival Differential LiDAR
abstract
Single-photon avalanche diode (SPAD) based LiDAR is becoming the de-facto choice for 3D imaging in many emerging applications. However, they suffer from three significant limitations: (a) the additional time-of-arrival dimension results in a data throughput bottleneck, (b) limited spatial resolution due to either low fill-factor (flash LiDAR) or scanning time (scanning-based LiDAR), and (c) coarse depth resolution due to quantization of photon timing by existing SPAD timing circuitries. In this paper, we present a novel, in-pixel computing architecture that we term first arrival differential (FAD) LiDAR, where instead of recording quantized time-of-arrival information at individual pixels, we record a temporal differential measurement between pairs of pixels. FAD captures relative order of photon arrivals at the two pixels (within a cycle or laser period) and creates a one-to-one mapping between this differential measurement and depth differences between the two pixels. We perform detailed system analysis and characterization using Monte Carlo simulation, and experimental emulation using a scanning-based single-photon avalanche diode. FAD pixels can result in a 10–100x reduction in per-pixel data throughput compared to TDC-based pixels. Under the same bandwidth constraints, FAD-LiDAR achieves better depth resolution and/or range than several state-of-the-art TDC-based LiDAR baselines.
Mel J. White, Akshat Dave, Shahaboddin Ghajari, Ankit Raghuram, Alyosha C. Molnar, Ashok Veeraraghavan
ICCP3
2022 A Differential SPAD Array Architecture in 0.18 μm CMOS for HDR Imaging
abstract
We propose a scalable architecture for a differential single-photon avalanche diode (D-SPAD) array which generates measurements based on differential time-of-arrival in lieu of absolute time-of-arrival. This design addresses the throughput bottleneck in conventional sensors that record time-of-arrival statistics (such as histograms or raw arrival-time data) directly. In addition, this design also mitigates saturation at the pixel level and at the counter, making it an ideal candidate for use when the scene being imaged covers a high dynamic range (HDR). The differential nature of the data also obviates the need for large digital circuitry such as high bit depth counters or time to digital converters (TDCs). A prototype test structure of 16 pixels was fabricated in 0.18 $\mu$m CMOS, and we show images reconstructed from this chip that illustrate its capabilities.
Mel J. White, Shahaboddin Ghajari, Akshat Dave, Ashok Veeraraghavan, Alyosha C. Molnar
ISCAS4
2019 SNLOS: Non-line-of-sight Scanning through Temporal Focusing
abstract
Over the last decade, several techniques have been developed for looking around the corner by exploiting the round-trip travel time of photons. Typically, these techniques necessitate the collection of a large number of measurements with varying virtual source and virtual detector locations. This data is then processed by a reconstruction algorithm to estimate the hidden scene. As a consequence, even when the region of interest in the hidden volume is small and limited, the acquisition time needed is large as the entire dataset has to be acquired and then processed.In this paper, we present the first example of scanning based non-line-of-sight imaging technique. The key idea is that if the virtual sources (pulsed sources) on the wall are delayed using a quadratic delay profile (much like the quadratic phase of a focusing lens), then these pulses arrive at the same instant at a single point in the hidden volume – the point being scanned. On the imaging side, applying quadratic delays to the virtual detectors before integration on a single gated detector allows us to ‘focus’ and scan each point in the hidden volume. By changing the quadratic delay profiles, we can focus light at different points in the hidden volume. This provides the first example of scanning based non-line-of-sight imaging, allowing us to focus our measurements only in the region of interest. We derive the theoretical underpinnings of ‘temporal focusing’, show compelling simulations of performance analysis, build a hardware prototype system and demonstrate real results.
Adithya Kumar Pediredla, Akshat Dave, Ashok Veeraraghavan
ICCP2
2019 Convolutional Approximations to the General Non-Line-of-Sight Imaging Operator
abstract
Non-line-of-sight (NLOS) imaging aims to reconstruct scenes outside the field of view of an imaging system. A common approach is to measure the so-called light transients, which facilitates reconstructions through ellipsoidal tomography that involves solving a linear least-squares. Unfortunately, the corresponding linear operator is very high-dimensional and lacks structures that facilitate fast solvers, and so, the ensuing optimization is a computationally daunting task. We introduce a computationally tractable framework for solving the ellipsoidal tomography problem. Our main observation is that the Gram of the ellipsoidal tomography operator is convolutional, either exactly under certain idealized imaging conditions, or approximately in practice. This, in turn, allows us to obtain the ellipsoidal tomography solution by using efficient deconvolution procedures to solve a linear least-squares problem involving the Gram operator. The computational tractability of our approach also facilitates the use of various regularizers during the deconvolution procedure. We demonstrate the advantages of our framework in a variety of simulated and real experiments.
Byeongjoo Ahn, Akshat Dave, Ashok Veeraraghavan, Ioannis Gkioulekas, Aswin C. Sankaranarayanan
ICCV2
2017 Compressive image recovery using recurrent generative model
abstract
Reconstruction of signals from compressively sensed measurements is an ill-posed problem. In this paper, we leverage the recurrent generative model, RIDE, as an image prior for compressive image reconstruction. Recurrent networks can model long-range dependencies in images and hence can handle global multiplexing in compressive imaging. We perform MAP inference with RIDE using back-propagation to the inputs and projected gradient method. We propose an entropy thresholding based approach for preserving texture in images well. Our approach shows superior reconstructions compared to recent global reconstruction approaches like D-AMP and TVAL3 on both simulated and real data.
Akshat Dave, Anil Kumar Vadathya, Kaushik Mitra
ICIP1
2014 Improving Saliency Models by Predicting Human Fixation Patches
Rachit Dubey, Akshat Dave, Bernard Ghanem
ACCV (3)2
2012 Do humans fixate on interest points?
Akshat Dave, Rachit Dubey, Bernard Ghanem
ICPR1