VLDB 2026 Research / reviewers in the wild / expert
Vishnu Sarukkai
dblp:255/5601
· DBLP profile ↗
7ranked-venue papers
4as first author
6since 2021 · last 2025
0009-0006-9809-9994ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Language models and text generation · 55% Trustworthy machine learning · 16% Learning theory · 16% | |
| Computer graphics and multimedia
2 papers |
Computer animation and physical simulation · 70% Visual content generation and editing · 30% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
in-context learning |
0.9 | 1 | 2025 | Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making Tasks · NeurIPS 2025 |
Natural language and speech › Language models and text generation
LLM agents |
0.9 | 1 | 2025 | Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making Tasks · NeurIPS 2025 |
Computer animation and physical simulation
character animation |
0.9 | 1 | 2025 | Learning to Ball: Composing Policies for Long-Horizon Basketball Moves · ACM Trans. Graph. 2025 |
Computer animation and physical simulation
motion control |
0.9 | 1 | 2025 | Learning to Ball: Composing Policies for Long-Horizon Basketball Moves · ACM Trans. Graph. 2025 |
Visual content generation and editing › image-to-image translation
sketch-to-image generation |
0.8 | 1 | 2024 | Block and Detail: Scaffolding Sketch-to-Image Generation · UIST 2024 |
Machine learning › Learning theory › classification
classifier evaluation |
0.5 | 1 | 2021 | Low-Shot Validation: Active Importance Sampling for Estimating Classifier Performance on Rare Categories · ICCV 2021 |
Machine learning › Trustworthy machine learning
model validation |
0.5 | 1 | 2021 | Low-Shot Validation: Active Importance Sampling for Estimating Classifier Performance on Rare Categories · ICCV 2021 |
Program synthesis and code generation
code generation with language models |
0.3 | 1 | 2025 | Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making Tasks · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
controlnet |
0.2 | 1 | 2024 | Block and Detail: Scaffolding Sketch-to-Image Generation · UIST 2024 |
Machine learning › Generative modeling
diffusion model |
0.2 | 1 | 2024 | Block and Detail: Scaffolding Sketch-to-Image Generation · UIST 2024 |
Methods — techniques the papers use, named apart from their topics
trajectory bootstrapping · 1.7population-based training · 1.7in-context examples · 1.7user study · 1.5dataset generation · 1.5controlnet · 1.5soft router · 0.9skill chaining · 0.9reinforcement learning · 0.9mixture of experts · 0.9variance estimation · 0.5importance sampling · 0.5active sampling · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making TasksabstractImproving Large Language Model (LLM) agents for sequential decision-making tasks typically requires extensive task-specific knowledge engineering—custom prompts, curated examples, and specialized observation/action spaces. We investigate a different approach where agents automatically improve by learning from their own successful experiences without human intervention. Our method constructs and refines a database of self-generated trajectories that serve as in-context examples for future tasks. Even naive accumulation of successful trajectories yields substantial performance gains across three diverse benchmarks: ALFWorld (73\% to 89\%), Wordcraft (55\% to 64\%), and InterCode-SQL (75\% to 79\%). These improvements exceed those achieved by upgrading from gpt-4o-mini to gpt-4o and match the performance of allowing multiple attempts per task. We further enhance this approach with two innovations: database-level curation using population-based training to propagate high-performing example collections, and exemplar-level curation that selectively retains trajectories based on their empirical utility as in-context examples. With these enhancements, our method achieves 93\% success on ALFWorld—surpassing approaches that use more powerful LLMs and hand-crafted components. Our trajectory bootstrapping technique demonstrates that agents can autonomously improve through experience, offering a scalable alternative to labor-intensive knowledge engineering. Vishnu Sarukkai, Kayvon Fatahalian |
NeurIPS | 1 |
| 2025 | Learning to Ball: Composing Policies for Long-Horizon Basketball MovesabstractLearning a control policy for a multi-phase, long-horizon task, such as basketball maneuvers, remains challenging for reinforcement learning approaches due to the need for seamless policy composition and transitions between skills. A long-horizon task typically consists of distinct subtasks with well-defined goals, separated by transitional subtasks with unclear goals but critical to the success of the entire task. Existing methods like the mixture of experts and skill chaining struggle with tasks where individual policies do not share significant commonly explored states or lack well-defined initial and terminal states between different phases. In this paper, we introduce a novel policy integration framework to enable the composition of drastically different motor skills in multi-phase long-horizon tasks with ill-defined intermediate states. Based on that, we further introduce a high-level soft router to enable seamless and robust transitions between the subtasks. We evaluate our framework on a set of fundamental basketball skills and challenging transitions. Policies trained by our approach can effectively control the simulated character to interact with the ball and accomplish the long-horizon task specified by real-time user commands, without relying on ball trajectory references. Pei Xu 0005, Ruocheng Wang, Vishnu Sarukkai, Kayvon Fatahalian, Ioannis Karamouzas, Victor B. Zordan, C. Karen Liu |
ACM Trans. Graph. | 4 |
| 2024 | Block and Detail: Scaffolding Sketch-to-Image GenerationabstractWe introduce a novel sketch-to-image tool that aligns with the iterative refinement process of artists. Our tool lets users sketch blocking strokes to coarsely represent the placement and form of objects and detail strokes to refine their shape and silhouettes. We develop a two-pass algorithm for generating high-fidelity images from such sketches at any point in the iterative process. In the first pass we use a ControlNet to generate an image that strictly follows all the strokes (blocking and detail) and in the second pass we add variation by renoising regions surrounding blocking strokes. We also present a dataset generation scheme that, when used to train a ControlNet architecture, allows regions that do not contain strokes to be interpreted as not-yet-specified regions rather than empty space. We show that this partial-sketch-aware ControlNet can generate coherent elements from partial sketches that only contain a small number of strokes. The high-fidelity images produced by our approach serve as scaffolds that can help the user adjust the shape and proportions of objects or add additional elements to the composition. We demonstrate the effectiveness of our approach with a variety of examples and evaluative comparisons. Quantitatively, evaluative user feedback indicates that novice viewers prefer the quality of images from our algorithm over a baseline Scribble ControlNet for 84% of the pairs and found our images had less distortion in 81% of the pairs. Vishnu Sarukkai, Mia Tang, Maneesh Agrawala, Kayvon Fatahalian |
UIST | 1 |
| 2024 | Collage DiffusionabstractWe seek to give users precise control over diffusion-based image generation by modeling complex scenes as sequences of layers, which define the desired spatial arrangement and visual attributes of objects in the scene. Collage Diffusion harmonizes the input layers to make objects fit together—the key challenge involves minimizing changes in the positions and key visual attributes of the input layers while allowing other attributes to change in the harmonization process. We ensure that objects are generated in the correct locations by modifying text-image cross-attention with the layers’ alpha masks. We preserve key visual attributes of input layers by learning specialized text representations per layer and by extending prior diffusion-based control mechanisms to operate on layers. Layer input allows users to control the extent of image harmonization on a per-object basis, and users can even iteratively edit individual objects in generated images while keeping other objects fixed. By leveraging the rich information present in layer input, Collage Diffusion generates globally harmonized images that maintain desired object characteristics better than prior approaches. Vishnu Sarukkai, Linden Li, Arden Ma, Christopher Ré, Kayvon Fatahalian |
WACV | 1 |
| 2024 | Learning to Move Like Professional Counter-Strike PlayersabstractAbstract In multiplayer, first‐person shooter games like Counter‐Strike: Global Offensive (CS:GO), coordinated movement is a critical component of high‐level strategic play. However, the complexity of team coordination and the variety of conditions present in popular game maps make it impractical to author hand‐crafted movement policies for every scenario. We show that it is possible to take a data‐driven approach to creating human‐like movement controllers for CS:GO. We curate a team movement dataset comprising 123 hours of professional game play traces, and use this dataset to train a transformer‐based movement model that generates human‐like team movement for all players in a “Retakes” round of the game. Importantly, the movement prediction model is efficient. Performing inference for all players takes less than 0.5 ms per game step (amortized cost) on a single CPU core, making it plausible for use in commercial games today. Human evaluators assess that our model behaves more like humans than both commercially‐available bots and procedural movement controllers scripted by experts (16% to 59% higher by TrueSkill rating of “human‐like”). Using experiments involving in‐game bot vs. bot self‐play, we demonstrate that our model performs simple forms of teamwork, makes fewer common movement mistakes, and yields movement distributions, player lifetimes, and kill locations similar to those observed in professional CS:GO match play. David Durst, Feng Xie 0008, Vishnu Sarukkai, Brennan Shacklett, Iuri Frosio, Chen Tessler, Joohwan Kim, Carly Taylor, Gilbert Louis Bernstein, Sanjiban Choudhury, Pat Hanrahan, Kayvon Fatahalian |
Comput. Graph. Forum | 3 |
| 2021 | Low-Shot Validation: Active Importance Sampling for Estimating Classifier Performance on Rare CategoriesabstractFor machine learning models trained with limited labeled training data, validation stands to become the main bottleneck to reducing overall annotation costs. We propose a statistical validation algorithm that accurately estimates the F-score of binary classifiers for rare categories, where finding relevant examples to evaluate on is particularly challenging. Our key insight is that simultaneous calibration and importance sampling enables accurate estimates even in the low-sample regime (< 300 samples). Critically, we also derive an accurate single-trial estimator of the variance of our method and demonstrate that this estimator is empirically accurate at low sample counts, enabling a practitioner to know how well they can trust a given low-sample estimate. When validating state-of-the-art semi-supervised models on ImageNet and iNatural-ist2017, our method achieves the same estimates of model performance with up to 10× fewer labels than competing approaches. In particular, we can estimate model F1 scores with a variance of 0.005 using as few as 100 labels. Fait Poms, Vishnu Sarukkai, Ravi Teja Mullapudi, Nimit Sharad Sohoni, William R. Mark, Deva Ramanan, Kayvon Fatahalian |
ICCV | 2 |
| 2020 | Cloud Removal in Satellite Images Using Spatiotemporal Generative NetworksabstractSatellite images hold great promise for continuous environmental monitoring and earth observation. Occlusions cast by clouds, however, can severely limit coverage, making ground information extraction more difficult. Existing pipelines typically perform cloud removal with simple temporal composites and hand-crafted filters. In contrast, we cast the problem of cloud removal as a conditional image synthesis challenge, and we propose a trainable spatiotemporal generator network (STGAN) to remove clouds. We train our model on a new large-scale spatiotemporal dataset that we construct, containing 97640 image pairs covering all continents. We demonstrate experimentally that the proposed STGAN model outperforms standard models and can generate realistic cloud-free images with high PSNR and SSIM values across a variety of atmospheric conditions, leading to improved performance in downstream tasks such as land cover classification. Vishnu Sarukkai, Anirudh Jain, Burak Uzkent, Stefano Ermon |
WACV | 1 |