VLDB 2026 Research / reviewers in the wild / expert
Saurabh Saxena
dblp:15/8791
· DBLP profile ↗
15ranked-venue papers
2as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 1 first-author · 9 since 2021Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | High-Resolution Frame Interpolation with Patch-based Cascaded DiffusionabstractDespite the recent progress, existing frame interpolation methods still struggle with processing extremely high resolution input and handling challenging cases such as repetitive textures, thin objects, and large motion. To address these issues, we introduce a patch-based cascaded pixel diffusion model for high resolution frame interpolation, HiFI, that excels in these scenarios while achieving competitive performance on standard benchmarks. Cascades, which generate a series of images from low to high resolution, can help significantly with large or complex motion that require both global context for a coarse solution and detailed context for high resolution output. However, contrary to prior work on cascaded diffusion models which perform diffusion on increasingly large resolutions, we use a single model that always performs diffusion at the same resolution and upsamples by processing patches of the inputs and the prior solution. At inference time, this drastically reduces memory usage and allows a single model, solving both frame interpolation (base model’s task) and spatial up-sampling, saving training cost as well. HiFI excels at high-resolution images and complex repeated textures that require global context, achieving comparable or state-of-the-art performance on various benchmarks (Vimeo, Xiph, X-Test, and SEPE-8K). We further introduce a new dataset, LaMoR, that focuses on particularly challenging cases, and HiFI significantly outperforms other baselines. Junhwa Hur, Charles Herrmann, Saurabh Saxena, Janne Kontkanen, Wei-Sheng Lai, Michael Rubinstein, David J. Fleet, Deqing Sun |
AAAI | 3 |
| 2025 | RoMo: Robust Motion Segmentation Improves Structure from MotionabstractThere has been extensive progress in the reconstruction and generation of 4D scenes from monocular casually-captured video. While these tasks rely heavily on known camera poses, the problem of finding such poses using structure-from-motion (SfM) often depends on robustly separating static from dynamic parts of a video. The lack of a robust solution to this problem limits the performance of SfM camera-calibration pipelines. We propose a novel approach to video-based motion segmentation to identify the components of a scene that are moving w.r.t. a fixed world frame. Our simple but effective iterative method, RoMo, combines optical flow and epipolar cues with a pre-trained video segmentation model. It outperforms unsupervised baselines for motion segmentation as well as supervised baselines trained from synthetic data. More importantly, the combination of an off-the-shelf SfM pipeline with our segmentation masks establishes a new state-of-the-art on camera calibration for scenes with dynamic content, outperforming existing methods by a substantial margin. Lily Goli, Sara Sabour, Mark J. Matthews, Marcus A. Brubaker, Dmitry Lagun, Alec Jacobson, David J. Fleet, Saurabh Saxena, Andrea Tagliasacchi |
ICCV | 8 |
| 2025 | Controlling Space and Time with Diffusion ModelsabstractWe present 4DiM, a cascaded diffusion model for 4D novel view synthesis (NVS), supporting generation with arbitrary camera trajectories and timestamps, in natural scenes, conditioned on one or more images. With a novel architecture and sampling procedure, we enable training on a mixture of 3D (with camera pose), 4D (pose+time) and video (time but no pose) data, which greatly improves generalization to unseen images and camera pose trajectories over prior works which generally operate in limited domains (e.g., object centric).
4DiM is the first-ever NVS method with intuitive metric-scale camera pose control enabled by our novel calibration pipeline for structure-from-motion-posed data. Experiments demonstrate that 4DiM outperforms prior 3D NVS models both in terms of
image fidelity and pose alignment, while also enabling the generation of scene dynamics. 4DiM provides a general framework for a variety of tasks including single-image-to-3D, two-image-to-video (interpolation and extrapolation), and pose-conditioned video-to-video translation, which we illustrate qualitatively on a variety of scenes.
See https://4d-diffusion.github.io for video samples. Daniel Watson, Saurabh Saxena, Lala Li, Andrea Tagliasacchi, David J. Fleet |
ICLR | 2 |
| 2024 | NeRFiller: Completing Scenes via Generative 3D InpaintingabstractWe propose NeRFiller, an approach that completes missing portions of a 3D capture via generative 3D inpainting using off-the-shelf 2D visual generative models. Often parts of a captured 3D scene or object are missing due to mesh reconstruction failures or a lack of observations (e.g., contact regions, such as the bottom of objects, or hard-to-reach areas). We approach this challenging 3D inpainting problem by leveraging a 2D inpainting diffusion model. We identify a surprising behavior of these models, where they generate more 3D consistent inpaints when images form a 2 x 2 grid, and show how to generalize this behavior to more than four images. We then present an iterative framework to distill these inpainted regions into a single consistent 3D scene. In contrast to related works, we focus on completing scenes rather than deleting foreground objects, and our approach does not require tight 2D object masks or text. We compare our approach to relevant baselines adapted to our setting on a variety of scenes, where NeRFiller creates the most 3D consistent and plausible scene completions. Our project page is at https://ethanweber.melnerfiller. Ethan Weber, Aleksander Holynski, Varun Jampani, Saurabh Saxena, Noah Snavely, Abhishek Kar, Angjoo Kanazawa |
CVPR | 4 |
| 2023 | A Generalist Framework for Panoptic Segmentation of Images and VideosabstractPanoptic segmentation assigns semantic and instance ID labels to every pixel of an image. As permutations of instance IDs are also valid solutions, the task requires learning of high-dimensional one-to-many mapping. As a result, state-of-the-art approaches use customized architectures and task-specific loss functions. We formulate panoptic segmentation as a discrete data generation problem, without relying on inductive bias of the task. A diffusion model is proposed to model panoptic masks, with a simple architecture and generic loss function. By simply adding past predictions as a conditioning signal, our method is capable of modeling video (in a streaming setting) and thereby learns to track object instances automatically. With extensive experiments, we demonstrate that our simple approach can perform competitively to state-of-the-art specialist methods in similar settings.1 Ting Chen 0007, Lala Li, Saurabh Saxena, Geoffrey E. Hinton, David J. Fleet |
ICCV | 3 |
| 2023 | The Surprising Effectiveness of Diffusion Models for Optical Flow and Monocular Depth EstimationabstractDenoising diffusion probabilistic models have transformed image generation with their impressive fidelity and diversity.
We show that they also excel in estimating optical flow and monocular depth, surprisingly without task-specific architectures and loss functions that are predominant for these tasks.
Compared to the point estimates of conventional regression-based methods, diffusion models also enable Monte Carlo inference, e.g., capturing uncertainty and ambiguity in flow and depth.
With self-supervised pre-training, the combined use of synthetic and real data for supervised training, and technical innovations (infilling and step-unrolled denoising diffusion training) to handle noisy-incomplete training data, one can train state-of-the-art diffusion models for depth and optical flow estimation, with additional zero-shot coarse-to-fine refinement for high resolution estimates.
Extensive experiments focus on quantitative performance against benchmarks, ablations, and the model's ability to capture uncertainty and multimodality, and impute missing values. Our model obtains a state-of-the-art relative depth error of 0.074 on the indoor NYU benchmark and an Fl-all score of 3.26\% on the KITTI optical flow benchmark, about 25\% better than the best published method. Saurabh Saxena, Charles Herrmann, Junhwa Hur, Abhishek Kar, Mohammad Norouzi 0002, Deqing Sun, David J. Fleet |
NeurIPS | 1 |
| 2022 | Pix2seq: A Language Modeling Framework for Object Detection
Ting Chen 0007, Saurabh Saxena, Lala Li, David J. Fleet, Geoffrey E. Hinton |
ICLR | 2 |
| 2022 | A Class-C Injection-Locked Tripler with 48 dB Sub-Harmonic Suppression and 15 fs Additive RMS Jitter in 0.13μm BiCMOS ProcessabstractWe present a low phase noise 4.5-to-6.5 GHz injection-locked oscillator-based frequency tripler (ILT) from an ultra-low jitter 1.5-to-2.16 GHz clock source. Class-C biasing is employed in the digitally controlled LC oscillator (LC-DCO) and the injection circuit to simultaneously achieve low phase noise in the LC-DCO and improve the third harmonic injection strength. Using BJT in the injection transistors and DCO greatly improves the low-frequency phase noise performance of the ILT. Fabricated in a $0.13 \mu \mathrm{m}$ BiCMOS process, the ILT has a measured tuning range of 4.5-to-6.5 GHz, with a jitter tracking bandwidth of 25 MHz for sub-harmonic injection. The ILT demonstrates good sub-harmonic rejection ratios SHRR1and SHRR2as 48 dB and 58 dB, respectively. The ILT adds 15 fs additional rms jitter to an input clock source with rms jitter of 68 fs. Sonam Sadhukhan, Pranav Kumar, Arpan Thakkar, Apoorva Bhatia, Saurabh Saxena |
ISCAS | 5 |
| 2022 | A Unified Sequence Interface for Vision TasksabstractWhile language tasks are naturally expressed in a single, unified, modeling framework, i.e., generating sequences of tokens, this has not been the case in computer vision. As a result, there is a proliferation of distinct architectures and loss functions for different vision tasks. In this work we show that a diverse set of "core" computer vision tasks can also be unified if formulated in terms of a shared pixel-to-sequence interface. We focus on four tasks, namely, object detection, instance segmentation, keypoint detection, and image captioning, all with diverse types of outputs, e.g., bounding boxes or dense masks. Despite that, by formulating the output of each task as a sequence of discrete tokens with a unified interface, we show that one can train a neural network with a single model architecture and loss function on all these tasks, with no task-specific customization. To solve a specific task, we use a short prompt as task description, and the sequence output adapts to the prompt so it can produce task-specific output. We show that such a model can achieve competitive performance compared to well-established task-specific models. Ting Chen 0007, Saurabh Saxena, Lala Li, Tsung-Yi Lin, David J. Fleet, Geoffrey E. Hinton |
NeurIPS | 2 |
| 2022 | Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingabstractWe present Imagen, a text-to-image diffusion model with an unprecedented degree of photorealism and a deep level of language understanding. Imagen builds on the power of large transformer language models in understanding text and hinges on the strength of diffusion models in high-fidelity image generation. Our key discovery is that generic large language models (e.g., T5), pretrained on text-only corpora, are surprisingly effective at encoding text for image synthesis: increasing the size of the language model in Imagen boosts both sample fidelity and image-text alignment much more than increasing the size of the image diffusion model. Imagen achieves a new state-of-the-art FID score of 7.27 on the COCO dataset, without ever training on COCO, and human raters find Imagen samples to be on par with the COCO data itself in image-text alignment. To assess text-to-image models in greater depth, we introduce DrawBench, a comprehensive and challenging benchmark for text-to-image models. With DrawBench, we compare Imagen with recent methods including VQ-GAN+CLIP, Latent Diffusion Models, and DALL-E 2, and find that human raters prefer Imagen over other models in side-by-side comparisons, both in terms of sample quality and image-text alignment. Chitwan Saharia, Saurabh Saxena, Lala Li, Jay Whang, Remi Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J. Fleet, Mohammad Norouzi 0002 |
NeurIPS | 3 |
| 2020 | Non-Autoregressive Machine Translation with Latent AlignmentsabstractThis paper presents two strong methods, CTC and Imputer, for non-autoregressive machine translation that model latent alignments with dynamic programming.We revisit CTC for machine translation and demonstrate that a simple CTC model can achieve state-of-theart for single-step non-autoregressive machine translation, contrary to what prior work indicates.In addition, we adapt the Imputer model for non-autoregressive machine translation and demonstrate that Imputer with just 4 generation steps can match the performance of an autoregressive Transformer baseline.Our latent alignment models are simpler than many existing non-autoregressive translation baselines; for example, we do not require target length prediction or re-scoring with an autoregressive model.On the competitive WMT'14 En→De task, our CTC model achieves 25.7 BLEU with a single generation step, while Imputer achieves 27.5 BLEU with 2 generation steps, and 28.0 BLEU with 4 generation steps.This compares favourably to the autoregressive Transformer baseline at 27.8 BLEU. Chitwan Saharia, Saurabh Saxena, Mohammad Norouzi 0002 |
EMNLP (1) | 3 |
| 2020 | A 2.5-5GHz Injection-Locked Clock Multiplier with Embedded Phase Interpolator in 65nm CMOSabstractWe present a wide-range ring oscillator based injection-locked clock multiplier with an embedded phase interpolator. The pseudo-differential ring oscillator employs current-starved inverters with a tunable capacitor load. The inverter delays are controlled to change the output phase while retaining the output frequency and jitter. Designed in a 65nm CMOS process, the clock multiplier uses a 312.5MHz reference clock and generates a 2.5–5.0GHz output clock when simulated with RC-extracted parasitics. The output clock phase can be controlled with 1.5° − 2° accuracy. At 5GHz output, the clock multiplier achieves an integrated rms jitter of 450–550fs across phase interpolation codes. Operating from 1.2V supply, it consumes 9.4mW at 5GHz that scales down to 4.3mW at 2.5GHz. R. Gautam, Jaya Deepthi Bandarupalli, Saurabh Saxena |
ISCAS | 3 |
| 2018 | Area and Current Efficient Capacitor-Less Low Drop-Out Regulator Using Time-Based Error AmplifierabstractAn output capacitor-less low drop-out (LDO) regulator using time-based error amplifier is presented in this paper. The proposed LDO utilizes voltage-controlled oscillator (VCO) as an integrator to replace the conventional voltage-based error amplifier. It reduces the overall area by using only 1.2pF of on-chip capacitor while consuming low quiescent current (<;30μA at typical corner). Using time as the processing variable, the time-based error amplifier operates with full-swing CMOS digital like signals without introducing any quantization error. The proposed LDO was designed in TSMC 65nm CMOS LP technology with input and output voltages as 1.2 V and 0.8V-1.1V, respectively, achieving a regulation bandwidth of 3MHz. Settling time of 200ns or less was achieved for 10mA load current step and 0-100pF output load capacitor. Qadeer A. Khan, Saurabh Saxena, Abirmoya Santra |
ISCAS | 2 |
| 2017 | Large-Scale Evolution of Image ClassifiersabstractNeural networks have proven effective at solving difficult problems but designing their architectures can be challenging, even for image classification problems alone. Our goal is to minimize human participation, so we employ evolutionary algorithms to discover such networks automatically. Despite significant computational requirements, we show that it is now possible to evolve models with accuracies within the range of those published in the last year. Specifically, we employ simple evolutionary techniques at unprecedented scales to discover models for the CIFAR-10 and CIFAR-100 datasets, starting from trivial initial conditions and reaching accuracies of 94.6\% (95.6\% for ensemble) and 77.0\%, respectively. To do this, we use novel and intuitive mutation operators that navigate large search spaces; we stress that no human participation is required once evolution starts and that the output is a fully-trained model. Throughout this work, we place special emphasis on the repeatability of results, the variability in the outcomes and the computational requirements. Esteban Real, Sherry Moore, Andrew Selle, Saurabh Saxena, Yutaka I. Leon-Suematsu, Jie Tan 0001, Quoc V. Le, Alexey Kurakin |
ICML | 4 |
| 2009 | Automatic Tuning of Time Constants in Single Bit Continuous-time Delta-sigma ModulatorsabstractWe describe an in-situ analog technique for estimating time constant shifts in continuous-time single bit delta-sigma modulators. We show that the variance of the first integrator output of the modulator's loop filter is a good indicator of RC time constants. The nominal values of the time constants are restored by digitally controlling the resistance and capacitance values (realized as switched banks). Simulation results that demonstrate the efficacy of the time-constant tuning system are given for a third order CIFF modulator. Saurabh Saxena, Prabu Sankar, Shanthi Pavan |
ISCAS | 1 |