Meera Hahn

dblp:173/5203 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
7since 2021 · last 2025
0009-0004-1549-088XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Generative modeling · 45% Deep learning architectures and training · 14% Robot navigation and mapping · 13%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › diffusion model
text-to-image generation
1.622025
Proactive Agents for Multi-Turn Text-to-Image Generation Under Uncertainty · ICML 2025
FineStyle: Fine-grained Controllable Style Personalization for Text-to-image Models · NeurIPS 2024
Machine learning › Generative modeling
video generation
1.522024
VideoPoet: A Large Language Model for Zero-Shot Video Generation · ICML 2024
Photorealistic Video Generation with Diffusion Models · ECCV (79) 2024
Natural language and speech › Language models and text generation › prompt tuning
prompt alignment
0.912025
Proactive Agents for Multi-Turn Text-to-Image Generation Under Uncertainty · ICML 2025
Human-AI interaction › AI agent
proactive agents
0.912025
Proactive Agents for Multi-Turn Text-to-Image Generation Under Uncertainty · ICML 2025
Machine learning › Generative modeling
diffusion model
0.812024
FineStyle: Fine-grained Controllable Style Personalization for Text-to-image Models · NeurIPS 2024
Machine learning › Transfer learning and domain adaptation › few-shot learning
few-shot adaptation
0.812024
FineStyle: Fine-grained Controllable Style Personalization for Text-to-image Models · NeurIPS 2024
Computer vision › Vision and language › vision-language model
multimodal large language model
0.812024
VideoPoet: A Large Language Model for Zero-Shot Video Generation · ICML 2024
Machine learning › Deep learning architectures and training
transformer
0.812024
VideoPoet: A Large Language Model for Zero-Shot Video Generation · ICML 2024
Machine learning › Deep learning architectures and training › transformer
transformer decoder
0.812024
VideoPoet: A Large Language Model for Zero-Shot Video Generation · ICML 2024
Machine learning › Generative modeling › video generation
zero-shot video generation
0.812024
VideoPoet: A Large Language Model for Zero-Shot Video Generation · ICML 2024
Robotics › Robot navigation and mapping › visual navigation
image-goal navigation
0.512021
No RL, No Simulation: Learning to Navigate without Navigating · NeurIPS 2021
Robotics › Robot navigation and mapping › learning-based navigation
self-supervised navigation
0.512021
No RL, No Simulation: Learning to Navigate without Navigating · NeurIPS 2021
Robotics › Robot navigation and mapping › localization › multi-robot localization
cooperative localization
0.412020
Where Are You? Localization from Embodied Dialog · EMNLP (1) 2020
Knowledge, reasoning and agents › Multi-agent systems › human-agent interaction
embodied conversational agents
0.412020
Where Are You? Localization from Embodied Dialog · EMNLP (1) 2020
Machine learning › Generative modeling › diffusion model › diffusion model training
simulation-free training
0.112021
No RL, No Simulation: Learning to Navigate without Navigating · NeurIPS 2021
Natural language and speech › Question answering and dialogue systems
visual dialog
0.112020
Where Are You? Localization from Embodied Dialog · EMNLP (1) 2020

Methods — techniques the papers use, named apart from their topics

clarification questions · 1.7belief graph · 1.7VQAScore · 1.7task-specific adaptation · 0.8pre-training · 0.8parameter-efficient fine-tuning · 0.8diffusion model · 0.8cross-attention · 0.8autoregressive transformer · 0.8adapter tuning · 0.8
YearPublicationVenuePosition
2025 Enabling Controllable, Identity Preserving, Non-Rigid Edits in Human-Centric Images
abstract
We approach the problem of inserting a person into a novel scene and controlling their pose via text guidance. Given an image of a person, a masked image of a scene, and a text description of the target pose, our model generates realistic, highly controllable images. We validate the robustness of our model’s true-to-text accuracy and identity preservation via a user study on in-the-wild images. In addition, we present a novel dataset containing pairs of frames from human-centric and action-rich videos, with text captions of the difference in human pose between frames. We also explore the challenges of controllable identity preservation for in-the-wild scenes and the failure modes of similar models. Our methods achieve a 10% increase in pose adherence ([email protected]) over comparable methods without compromising visual fidelity, and show a clear qualitative improvement.
Nikolai Warner, Jack Kolb, Meera Hahn, Jonathan Huang, Vighnesh Birodkar, Irfan A. Essa
ICIP3
2025 Proactive Agents for Multi-Turn Text-to-Image Generation Under Uncertainty
abstract
User prompts for generative AI models are often underspecified, leading to a misalignment between the user intent and models’ understanding. As a result, users commonly have to painstakingly refine their prompts. We study this alignment problem in text-to-image (T2I) generation and propose a prototype for proactive T2I agents equipped with an interface to (1) actively ask clarification questions when uncertain, and (2) present their uncertainty about user intent as an understandable and editable belief graph. We build simple prototypes for such agents and propose a new scalable and automated evaluation approach using two agents, one with a ground truth intent (an image) while the other tries to ask as few questions as possible to align with the ground truth. We experiment over three image-text datasets: ImageInWords (Garg et al., 2024), COCO (Lin et al., 2014) and DesignBench, a benchmark we curated with strong artistic and design elements. Experiments over the three datasets demonstrate the proposed T2I agents’ ability to ask informative questions and elicit crucial information to achieve successful alignment with at least 2 times higher VQAScore (Lin et al., 2024) than the standard T2I generation. Moreover, we conducted human studies and observed that at least 90% of human subjects found these agents and their belief graphs helpful for their T2I workflow, highlighting the effectiveness of our approach. Code and DesignBench can be found at https://github.com/google-deepmind/proactive_t2i_agents.
Meera Hahn, Nithish Kannen, Rich Galt, Kartikeya Badola, Been Kim
ICML1
2024 Photorealistic Video Generation with Diffusion Models
Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu, Meera Hahn, Li Fei-Fei 0001, Irfan A. Essa, Lu Jiang 0004, José Lezama
ECCV (79)5
2024 VideoPoet: A Large Language Model for Zero-Shot Video Generation
abstract
We present VideoPoet, a language model capable of synthesizing high-quality video from a large variety of conditioning signals. VideoPoet employs a decoder-only transformer architecture that processes multimodal inputs – including images, videos, text, and audio. The training protocol follows that of Large Language Models (LLMs), consisting of two stages: pretraining and task-specific adaptation. During pretraining, VideoPoet incorporates a mixture of multimodal generative objectives within an autoregressive Transformer framework. The pretrained LLM serves as a foundation that can be adapted for a range of video generation tasks. We present empirical results demonstrating the model’s state-of-the-art capabilities in zero-shot video generation, specifically highlighting the ability to generate high-fidelity motions. Project page: http://sites.research.google/videopoet/
Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, Krishna Somandepalli, Hassan Akbari, Yair Alon, Yong Cheng 0003, Joshua V. Dillon, Agrim Gupta, Meera Hahn, Anja Hauth, David Hendon, Alonso Martinez, David Minnen, Mikhail Sirotenko, Kihyuk Sohn, Hartwig Adam, Ming-Hsuan Yang 0001, Irfan A. Essa, Huisheng Wang, David A. Ross, Bryan Seybold, Lu Jiang 0004
ICML17
2024 FineStyle: Fine-grained Controllable Style Personalization for Text-to-image Models
abstract
Few-shot fine-tuning of text-to-image (T2I) generation models enables people to create unique images in their own style using natural languages without requiring extensive prompt engineering. However, fine-tuning with only a handful, as little as one, of image-text paired data prevents fine-grained control of style attributes at generation. In this paper, we present FineStyle, a few-shot fine-tuning method that allows enhanced controllability for style personalized text-to-image generation. To overcome the lack of training data for fine-tuning, we propose a novel concept-oriented data scaling that amplifies the number of image-text pair, each of which focuses on different concepts (e.g., objects) in the style reference image. We also identify the benefit of parameter-efficient adapter tuning of key and value kernels of cross-attention layers. Extensive experiments show the effectiveness of FineStyle at following fine-grained text prompts and delivering visual quality faithful to the specified style, measured by CLIP scores and human raters.
Gong Zhang 0011, Kihyuk Sohn, Meera Hahn, Humphrey Shi, Irfan A. Essa
NeurIPS3
2023 Text and Click inputs for unambiguous open vocabulary instance segmentation
Vighnesh Birodkar, Jonathan Huang, Meera Hahn, Irfan A. Essa, Nikolai Warner
BMVC3
2021 No RL, No Simulation: Learning to Navigate without Navigating
abstract
Most prior methods for learning navigation policies require access to simulation environments, as they need online policy interaction and rely on ground-truth maps for rewards. However, building simulators is expensive (requires manual effort for each and every scene) and creates challenges in transferring learned policies to robotic platforms in the real-world, due to the sim-to-real domain gap. In this paper, we pose a simple question: Do we really need active interaction, ground-truth maps or even reinforcement-learning (RL) in order to solve the image-goal navigation task? We propose a self-supervised approach to learn to navigate from only passive videos of roaming. Our approach, No RL, No Simulator (NRNS), is simple and scalable, yet highly effective. NRNS outperforms RL-based formulations by a significant margin. We present NRNS as a strong baseline for any future image-based navigation tasks that use RL or Simulation.
Meera Hahn, Devendra Singh Chaplot, Shubham Tulsiani, Mustafa Mukadam, James M. Rehg, Abhinav Gupta 0001
NeurIPS1
2020 Tripping through time: Efficient Localization of Activities in Videos
Meera Hahn, Asim Kadav, James M. Rehg, Hans Peter Graf
BMVC1
2020 Where Are You? Localization from Embodied Dialog
abstract
We present WHERE ARE YOU? (WAY), a dataset of ~6k dialogs in which two humans – an Observer and a Locator – complete a cooperative localization task. The Observer is spawned at random in a 3D environment and can navigate from first-person views while answering questions from the Locator. The Locator must localize the Observer in a detailed top-down map by asking questions and giving instructions. Based on this dataset, we define three challenging tasks: Localization from Embodied Dialog or LED (localizing the Observer from dialog history), Embodied Visual Dialog (modeling the Observer), and Cooperative Localization (modeling both agents). In this paper, we focus on the LED task – providing a strong baseline model with detailed ablations characterizing both dataset biases and the importance of various modeling choices. Our best model achieves 32.7% success at identifying the Observer's location within 3m in unseen buildings, vs. 70.4% for human Locators.
Meera Hahn, Jacob Krantz, Dhruv Batra, Devi Parikh, James M. Rehg, Stefan Lee
EMNLP (1)1
2017 Situated Bayesian Reasoning Framework for Robots Operating in Diverse Everyday Environments
Sonia Chernova, Vivian Chu, Angel Andres Daruna, Haley Garrison, Meera Hahn, Priyanka Khante, Andrea Thomaz
ISRR5