Rohun Agrawal

dblp:332/1848 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Vision and language · 44% Knowledge representation and reasoning · 44% 3D vision · 13%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Knowledge representation and reasoning › spatial reasoning
3d spatial reasoning
0.912025
Visual Agentic AI for Spatial Reasoning with a Dynamic API · CVPR 2025
Computer vision › Vision and language
visual reasoning
0.912025
Visual Agentic AI for Spatial Reasoning with a Dynamic API · CVPR 2025
Program synthesis and code generation › neural program synthesis
LLM-based program synthesis
0.912025
Visual Agentic AI for Spatial Reasoning with a Dynamic API · CVPR 2025
Computer vision › 3D vision
3d scene understanding
0.312025
Visual Agentic AI for Spatial Reasoning with a Dynamic API · CVPR 2025

Methods — techniques the papers use, named apart from their topics

program synthesis · 1.7large language model agents · 0.9large language model agent · 0.9
YearPublicationVenuePosition
2025 Visual Agentic AI for Spatial Reasoning with a Dynamic API
abstract
Visual reasoning – the ability to interpret the visual world–is crucial for embodied agents that operate within three-dimensional scenes. Progress in AI has led to vision and language models capable of answering questions from images. However, their performance declines when tasked with 3D spatial reasoning. To tackle the complexity of such reasoning problems, we introduce an agentic program synthesis approach where LLM agents collaboratively generate a Pythonic API with new functions to solve common subproblems. Our method overcomes limitations of prior approaches that rely on a static, human-defined API, allowing it to handle a wider range of queries. To assess AI capabilities for 3D understanding, we introduce a new benchmark of queries involving multiple steps of grounding and inference. We show that our method outperforms prior zero-shot models for visual reasoning in 3D and empirically validate the effectiveness of our agentic framework for 3D spatial reasoning tasks. Project website: https://glab-caltech.github.io/vadar/
Damiano Marsili, Rohun Agrawal, Yisong Yue, Georgia Gkioxari
CVPR2
2023 Alternating Phase Langevin Sampling with Implicit Denoiser Priors for Phase Retrieval
abstract
Phase retrieval is the nonlinear inverse problem of recovering a true signal from its Fourier magnitude measurements. It arises in many applications such as astronomical imaging, X-Ray crystallography, microscopy, and more. The problem is highly ill-posed due to the phase-induced ambiguities and the large number of possible images that can fit to the given measurements. Thus, there’s a rich history of enforcing structural priors to improve solutions including sparsity priors and deep-learning-based generative models. However, such priors are often limited in their representational capacity or generalizability to slightly different distributions. Recent advancements in using denoisers as regularizers for non-convex optimization algorithms have shown promising performance and generalization. We present a way of leveraging the prior implicitly learned by a denoiser to solve phase retrieval problems by incorporating it in a classical alternating minimization framework. Compared to performant denoising-based algorithms for phase retrieval, we showcase competitive performance with Fourier measurements on in-distribution images and notable improvement on out-of-distribution images.
Rohun Agrawal, Oscar Leong
ICASSP1