Sakshum Kulshrestha

dblp:342/7791 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
0009-0005-8125-8867ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Autonomous driving · 42% Generative modeling · 41% 3D vision · 17%
Computer graphics and multimedia
3 papers
Computational photography and imaging · 66% Image and video processing · 34%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.622025
SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model · CVPR 2025
SceneDiffuser: Efficient and Controllable Driving Simulation Initialization and Rollout · NeurIPS 2024
Robotics › Autonomous driving › simulation
traffic simulation
1.622025
SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model · CVPR 2025
SceneDiffuser: Efficient and Controllable Driving Simulation Initialization and Rollout · NeurIPS 2024
Computational photography and imaging › computational optics
point-spread-function engineering
1.422024
CodedEvents: Optimal Point-Spread-Function Engineering for 3D-Tracking with Event Cameras · CVPR 2024
TiDy-PSFs: Computational Imaging with Time-Averaged Dynamic Point-Spread-Functions · ICCV 2023
Image and video processing › image restoration
image denoising
0.812024
A Scalable Training Strategy for Blind Multi-Distribution Noise Removal · IEEE Trans. Image Process. 2024
Image and video processing
image restoration
0.812024
A Scalable Training Strategy for Blind Multi-Distribution Noise Removal · IEEE Trans. Image Process. 2024
Computational photography and imaging
depth estimation
0.712023
TiDy-PSFs: Computational Imaging with Time-Averaged Dynamic Point-Spread-Functions · ICCV 2023
Computational photography and imaging › depth estimation
monocular depth estimation
0.712023
TiDy-PSFs: Computational Imaging with Time-Averaged Dynamic Point-Spread-Functions · ICCV 2023
Robotics › Autonomous driving
trajectory prediction
0.312025
SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model · CVPR 2025
Machine learning › Generative modeling
scene generation
0.212024
SceneDiffuser: Efficient and Controllable Driving Simulation Initialization and Rollout · NeurIPS 2024
Computational photography and imaging › depth of field
extended depth of field
0.212023
TiDy-PSFs: Computational Imaging with Time-Averaged Dynamic Point-Spread-Functions · ICCV 2023

Methods — techniques the papers use, named apart from their topics

phase mask design · 1.5implicit neural representation · 1.5amplitude mask design · 1.5scene generation · 0.9diffusion model · 0.9agent behavior modeling · 0.9polynomial approximation · 0.8inference-time constraint · 0.8diffusion · 0.8adaptive sampling · 0.8active learning · 0.8time-averaged PSF · 0.7spatial light modulator · 0.7deep learning · 0.7
YearPublicationVenuePosition
2025 SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model
abstract
The goal of traffic simulation is to augment a potentially limited amount of manually-driven miles that is available for testing and validation, with a much larger amount of simulated synthetic miles. The culmination of this vision would be a generative simulated city, where given a map of the city and an autonomous vehicle (AV) software stack, the simulator can seamlessly simulate the trip from point A to point B by populating the city around the AV and controlling all aspects of the scene, from animating the dynamic agents (e.g., vehicles, pedestrians) to controlling the traffic light states. We refer to this vision as CitySim, which requires an agglomeration of simulation technologies: scene generation to populate the initial scene, agent behavior modeling to animate the scene, occlusion reasoning, dynamic scene generation to seamlessly spawn and remove agents, and environment simulation for factors such as traffic lights. While some key technologies have been separately studied in various works, others such as dynamic scene generation and environment simulation have received less attention in the research community. We propose SceneDiffuser++, the first end-to-end generative world model trained on a single loss function capable of point A-to-B simulation on a city scale integrating all the requirements above. We demonstrate the city-scale traffic simulation capability of SceneDiffuser++ and study its superior realism under long simulation conditions. We evaluate the simulation quality on an augmented version of the Waymo Open Motion Dataset (WOMD) with larger map regions to support trip-level simulation.
Shuhan Tan, John Lambert, Hong Jeon, Sakshum Kulshrestha, Yijing Bai, Dragomir Anguelov, Mingxing Tan, Chiyu Max Jiang
CVPR4
2024 CodedEvents: Optimal Point-Spread-Function Engineering for 3D-Tracking with Event Cameras
abstract
Point-spread-function (PSF) engineering is a well-established computational imaging technique that uses phase masks and other optical elements to embed extra information (e.g., depth) into the images captured by conventional CMOS image sensors. To date, however, PSF-engineering has not been applied to neuromorphic event cameras; a powerful new image sensing technology that responds to changes in the log-intensity of light. This paper establishes theoretical limits (Cramér Rao bounds) on 3D point localization and tracking with PSF-engineered event cameras. Using these bounds, we first demonstrate that existing Fisher phase masks are already near-optimal for localizing static flashing point sources (e.g., blinking fluorescent molecules). We then demonstrate that existing designs are sub-optimal for tracking moving point sources and proceed to use our theory to design optimal phase masks and binary amplitude masks for this task. To overcome the non-convexity of the design problem, we leverage novel implicit neural representation based parameterizations of the phase and amplitude masks. We demonstrate the efficacy of our designs through extensive simulations. We also validate our method with a simple prototype.
Sachin Shah, Matthew A. Chan 0002, Haoming Cai, Jingxi Chen, Sakshum Kulshrestha, Chahat Deep Singh, Yiannis Aloimonos, Christopher A. Metzler
CVPR5
2024 SceneDiffuser: Efficient and Controllable Driving Simulation Initialization and Rollout
abstract
Simulation with realistic and interactive agents represents a key task for autonomous vehicle (AV) software development in order to test AV performance in prescribed, often long-tail scenarios. In this work, we propose SceneDiffuser, a scene-level diffusion prior for traffic simulation. We present a singular framework that unifies two key stages of simulation: scene initialization and scene rollout. Scene initialization refers to generating the initial layout for the traffic in a scene, and scene rollout refers to closed-loop simulation for the behaviors of the agents. While diffusion has been demonstrated to be effective in learning realistic, multimodal agent distributions, two open challenges remain: controllability and closed-loop inference efficiency and realism. To this end, to address controllability challenges, we propose generalized hard constraints, a generalized inference-time constraint mechanism that is simple yet effective. To improve closed-loop inference quality and efficiency, we propose amortized diffusion, a novel diffusion denoising paradigm that amortizes the physical cost of denoising over future simulation rollout steps, reducing the cost of per physical rollout step to a single denoising function evaluation, while dramatically reducing closed-loop errors. We demonstrate the effectiveness of our approach on the Waymo Open Dataset, where we are able to generate distributionally realistic scenes, while obtaining competitive performance in the Sim Agents Challenge, surpassing the state-of-the-art in many realism attributes.
Chiyu Max Jiang, Yijing Bai, Andre Cornman, Xiukun Huang, Hong Jeon, Sakshum Kulshrestha, John Lambert, Shuangyu Li, Xuanyu Zhou, Carlos Fuertes, Chang Yuan, Mingxing Tan, Dragomir Anguelov
NeurIPS7
2024 A Scalable Training Strategy for Blind Multi-Distribution Noise Removal
abstract
Despite recent advances, developing general-purpose universal denoising and artifact-removal networks remains largely an open problem: Given fixed network weights, one inherently trades-off specialization at one task (e.g., removing Poisson noise) for performance at another (e.g., removing speckle noise). In addition, training such a network is challenging due to the curse of dimensionality: As one increases the dimensions of the specification-space (i.e., the number of parameters needed to describe the noise distribution) the number of unique specifications one needs to train for grows exponentially. Uniformly sampling this space will result in a network that does well at very challenging problem specifications but poorly at easy problem specifications, where even large errors will have a small effect on the overall mean squared error. In this work we propose training denoising networks using an adaptive-sampling/active-learning strategy. Our work improves upon a recently proposed universal denoiser training strategy by extending these results to higher dimensions and by incorporating a polynomial approximation of the true specification-loss landscape. This approximation allows us to reduce training times by almost two orders of magnitude. We test our method on simulated joint Poisson-Gaussian-Speckle noise and demonstrate that with our proposed training strategy, a single blind, generalist denoiser network can achieve peak signal-to-noise ratios within a uniform bound of specialized denoiser networks across a large range of operating conditions. We also capture a small dataset of images with varying amounts of joint Poisson-Gaussian-Speckle noise and demonstrate that a universal denoiser trained using our adaptive-sampling strategy outperforms uniformly trained baselines.
Kevin Zhang 0003, Sakshum Kulshrestha, Christopher A. Metzler
IEEE Trans. Image Process.2
2023 TiDy-PSFs: Computational Imaging with Time-Averaged Dynamic Point-Spread-Functions
abstract
Point-spread-function (PSF) engineering is a powerful computational imaging technique wherein a custom phase mask is integrated into an optical system to encode additional information into captured images. Used in combination with deep learning, such systems now offer state-of-the-art performance at monocular depth estimation, extended depth-of-field imaging, lensless imaging, and other tasks. Inspired by recent advances in spatial light modulator (SLM) technology, this paper answers a natural question: Can one encode additional information and achieve superior performance by changing a phase mask dynamically over time? We first prove that the set of PSFs described by static phase masks is non-convex and that, as a result, time-averaged PSFs generated by dynamic phase masks are fundamentally more expressive. We then demonstrate, in simulation, that time-averaged dynamic (TiDy) phase masks can leverage this increased expressiveness to offer substantially improved monocular depth estimation and extended depth-of-field imaging performance.
Sachin Shah, Sakshum Kulshrestha, Christopher A. Metzler
ICCV2