EDBT 2026 Demo / reviewers in the wild / expert
Sanjay Subramanian
dblp:228/8258
· DBLP profile ↗
13ranked-venue papers
3as first author
10since 2021 · last 2025
0009-0001-1399-9920ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Vision and language · 43% Language models and text generation · 14% Trustworthy machine learning · 7% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 54% Programming languages and type systems · 46% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% |
Topics — the 19 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
visual reasoning |
1.4 | 2 | 2024 | Recursive Visual Programming · ECCV (43) 2024 From Wrong To Right: A Recursive Approach Towards Vision-Language Explanation · EMNLP 2023 |
Computer vision › Face, body and person analysis
human pose estimation |
0.9 | 1 | 2025 | Pose Priors from Language Models · CVPR 2025 |
Computer vision › 3D vision › geometric optimization
pose optimization |
0.9 | 1 | 2025 | Pose Priors from Language Models · CVPR 2025 |
Visual content generation and editing › video authoring
slideshow creation |
0.9 | 1 | 2025 | AutoPresent: Designing Structured Visuals from Scratch · CVPR 2025 |
Compilers and program optimization
code generation |
0.9 | 1 | 2025 | AutoPresent: Designing Structured Visuals from Scratch · CVPR 2025 |
Natural language and speech › Language models and text generation
large language model |
0.8 | 1 | 2024 | Using Language Models to Disambiguate Lexical Choices in Translation · EMNLP 2024 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
multi-agent planning |
0.8 | 1 | 2024 | TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering · EMNLP 2024 |
Computer vision › Video understanding and tracking
video question answering |
0.8 | 1 | 2024 | TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering · EMNLP 2024 |
Programming languages and type systems › programming paradigms
visual programming |
0.8 | 1 | 2024 | Recursive Visual Programming · ECCV (43) 2024 |
Computer vision › Vision and language › vision-language generation
visual explanation generation |
0.7 | 1 | 2023 | From Wrong To Right: A Recursive Approach Towards Vision-Language Explanation · EMNLP 2023 |
Computer vision › Vision and language › visual grounding
referring expression comprehension |
0.6 | 1 | 2022 | ReCLIP: A Strong Zero-Shot Baseline for Referring Expression Comprehension · ACL (1) 2022 |
Computer vision › Vision and language › multimodal understanding
zero-shot vision-language understanding |
0.6 | 1 | 2022 | ReCLIP: A Strong Zero-Shot Baseline for Referring Expression Comprehension · ACL (1) 2022 |
Computer vision › Vision and language › multimodal reasoning
compositional reasoning |
0.4 | 1 | 2020 | Obtaining Faithful Interpretations from Compositional Neural Networks · ACL 2020 |
Machine learning › Trustworthy machine learning › interpretability › explanation evaluation
explanation faithfulness |
0.4 | 1 | 2020 | Obtaining Faithful Interpretations from Compositional Neural Networks · ACL 2020 |
Machine learning › Trustworthy machine learning
interpretability |
0.4 | 1 | 2020 | Obtaining Faithful Interpretations from Compositional Neural Networks · ACL 2020 |
Computer vision › Vision and language
neural module network |
0.4 | 1 | 2020 | Obtaining Faithful Interpretations from Compositional Neural Networks · ACL 2020 |
Natural language and speech › Information extraction and text analysis › relation extraction › event relation extraction
temporal relation extraction |
0.4 | 1 | 2019 | An Improved Neural Baseline for Temporal Relation Extraction · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Language models and text generation
instruction following |
0.3 | 1 | 2025 | AutoPresent: Designing Structured Visuals from Scratch · CVPR 2025 |
Computer vision › Vision and language › vision-language model › multimodal large language model
video multimodal large language model |
0.2 | 1 | 2024 | TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering · EMNLP 2024 |
Methods — techniques the papers use, named apart from their topics
large language model · 3.4iterative self-refinement · 2.6large multimodal model · 1.6loss function · 0.9contact descriptor extraction · 0.9replanning · 0.8prompting · 0.8agent planning · 0.8recursive visual explanation · 0.7few-shot self-training · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AutoPresent: Designing Structured Visuals from ScratchabstractDesigning structured visuals such as presentation slides is essential for communicative needs, necessitating both content creation and visual planning skills. In this work, we tackle the challenge of automated slide generation, where models produce slide presentations from natural language (NL) instructions. We first introduce the SlidesBench benchmark, the first benchmark for slide generation with 7k training and 585 testing examples derived from 310 slide decks across 10 domains. SlidesBench supports evaluations that are (i) reference-based to measure similarity to a target slide, and (ii) reference-free to measure the design quality of generated slides alone. We benchmark end-to-end image generation and program generation methods with a variety of models, and find that programmatic methods produce higher-quality slides in user-interactable formats. Built on the success of program generation, we create AutoPresent, an 8B LlaMa-based model trained on 7k pairs of instructions paired with code for slide generation, and achieve results comparable to the closed-source model GPT-4O. We further explore iterative design refinement where the model is tasked to self-refine its own output, and we found that this process improves the slide’s quality. We hope that our work will provide a basis for future work on generating structured visuals. Our code, data, demo, and video demonstrations are publicly available at https: //github.com/para-lost/AutoPresent Jiaxin Ge, Zhiruo Wang 0001, Yi-Hao Peng, Sanjay Subramanian, Qinyue Tan, Maarten Sap, Alane Suhr, Daniel Fried, Graham Neubig, Trevor Darrell |
CVPR | 5 |
| 2025 | Pose Priors from Language ModelsabstractLanguage is often used to describe physical interaction, yet most 3D human pose estimation methods overlook this rich source of information. We bridge this gap by leveraging large multimodal models (LMMs) as priors for reconstructing contact poses, offering a scalable alternative to traditional methods that rely on human annotations or motion capture data. Our approach extracts contact-relevant descriptors from an LMM and translates them into tractable losses to constrain 3D human pose optimization. Despite its simplicity, our method produces compelling reconstructions for both two-person interactions and self-contact scenarios, accurately capturing the semantics of physical and social interactions. Our results demonstrate that LMMs can serve as powerful tools for contact prediction and pose estimation, offering an alternative to costly manual human annotations or motion capture data. Our code is publicly available at https://prosepose.github.io. Sanjay Subramanian, Evonne Ng, Lea Müller, Daniel Klein 0001, Shiry Ginosar, Trevor Darrell |
CVPR | 1 |
| 2024 | Recursive Visual Programming
Jiaxin Ge, Sanjay Subramanian, Baifeng Shi, Roei Herzig, Trevor Darrell |
ECCV (43) | 2 |
| 2024 | Using Language Models to Disambiguate Lexical Choices in TranslationabstractIn translation, a concept represented by a single word in a source language can have multiple variations in a target language.The task of lexical selection requires using context to identify which variation is most appropriate for a source text.We work with native speakers of nine languages to create DTAiLS, a dataset of 1,377 sentence pairs that exhibit cross-lingual concept variation when translating from English.We evaluate recent LLMs and neural machine translation systems on DTAiLS, with the bestperforming model, GPT-4, achieving from 67 to 85% accuracy across languages.Finally, we use language models to generate English rules describing target-language concept variations.Providing weaker models with high-quality lexical rules improves accuracy substantially, in some cases reaching or outperforming GPT-4. Concept: date (fruit)Variations and Generated Rules Khorma refers to the fruit of the date palm when it is fully ripe and dried.It is commonly consumed as a sweet, chewy snack or used in various dishes, particularly desserts.Rotab refers to fresh, soft dates that are partially ripe.These dates are moister and sweeter than fully ripe, dried dates (khorma).Rotab is often eaten as a fresh fruit or used in cooking where a softer, sweeter texture is desired.Kharak refers to dates that are unripe and are in a semi-dried state.They are less sweet compared to rotab and khorma and are often used in cooking or further processed into other forms. Josh Barua, Sanjay Subramanian, Kayo Yin, Alane Suhr |
EMNLP | 2 |
| 2024 | TraveLER: A Modular Multi-LMM Agent Framework for Video Question-AnsweringabstractRecently, image-based Large Multimodal Models (LMMs) have made significant progress in video question-answering (VideoQA) using a frame-wise approach by leveraging large-scale pretraining in a zero-shot manner. Nevertheless, these models need to be capable of finding relevant information, extracting it, and answering the question simultaneously. Currently, existing methods perform all of these steps in a single pass without being able to adapt if insufficient or incorrect information is collected. To overcome this, we introduce a modular multi-LMM agent framework based on several agents with different roles, instructed by a Planner agent that updates its instructions using shared feedback from the other agents. Specifically, we propose TraveLER, a method that can create a plan to "Traverse” through the video, ask questions about individual frames to "Locate” and store key information, and then "Evaluate” if there is enough information to answer the question. Finally, if there is not enough information, our method is able to "Replan” based on its collected knowledge. Through extensive experiments, we find that the proposed TraveLER approach improves performance on several VideoQA benchmarks without the need to fine-tune on specific datasets. Our code is available at https://github.com/traveler-framework/TraveLER. Chuyi Shang, Amos You, Sanjay Subramanian, Trevor Darrell, Roei Herzig |
EMNLP | 3 |
| 2023 | From Wrong To Right: A Recursive Approach Towards Vision-Language ExplanationabstractAddressing the challenge of adapting pretrained vision-language models for generating insightful explanations for visual reasoning tasks with limited annotations, we present ReVisE: a Recursive Visual Explanation algorithm.Our method iteratively computes visual features (conditioned on the text input), an answer, and an explanation, to improve the explanation quality step by step until the answer converges.We find that this multi-step approach guides the model to correct its own answers and outperforms single-step explanation generation.Furthermore, explanations generated by ReVisE also serve as valuable annotations for few-shot self-training.Our approach outperforms previous methods while utilizing merely 5% of the human-annotated explanations across 10 metrics, demonstrating up to a 4.2 and 1.3 increase in BLEU-1 score on the VCR and VQA-X datasets, underscoring the efficacy and dataefficiency of our method. Jiaxin Ge, Sanjay Subramanian, Trevor Darrell, Boyi Li 0001 |
EMNLP | 2 |
| 2023 | MEVade: An MEV-Resistant Blockchain DesignabstractEthereum is a popular blockchain that facilitates the creation of decentralized applications (dApps) and enables digital transactions to be executed without the need for a central authority. However, as in traditional markets, information asymmetry and market inefficiencies are used to the detriment of ordinary users via trading strategies that exploit “Miner Extractable Value” (MEV). We propose two extensions of Ethereum, one for proof of work (PoW), and one for proof of stake (PoS), that eliminate most forms of MEV by randomizing the execution order of transactions and hiding the content of transactions until their inclusion in a block. We simulate attack scenarios for both settings and provide detailed security properties and proofs. Julien Piet, Vivek Nair, Sanjay Subramanian |
ICBC | 3 |
| 2023 | Can Language Models Learn to Listen?abstractWe present a framework for generating appropriate facial responses from a listener in dyadic social interactions based on the speaker’s words. Given an input transcription of the speaker’s words with their timestamps, our approach autoregressively predicts a response of a listener: a sequence of listener facial gestures, quantized using a VQ-VAE. Since gesture is a language component, we propose treating the quantized atomic motion elements as additional language token inputs to a transformer-based large language model. Initializing our transformer with the weights of a language model pre-trained only on text results in significantly higher quality listener responses than training a transformer from scratch. We show that our generated listener motion is fluent and reflective of language semantics through quantitative metrics and a qualitative user study. In our evaluation, we analyze the model’s ability to utilize temporal and semantic aspects of spoken text. Evonne Ng, Sanjay Subramanian, Daniel Klein 0001, Angjoo Kanazawa, Trevor Darrell, Shiry Ginosar |
ICCV | 2 |
| 2022 | ReCLIP: A Strong Zero-Shot Baseline for Referring Expression ComprehensionabstractSanjay Subramanian, William Merrill, Trevor Darrell, Matt Gardner, Sameer Singh, Anna Rohrbach. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Sanjay Subramanian, William Merrill, Trevor Darrell, Matt Gardner 0001, Sameer Singh 0001, Anna Rohrbach |
ACL (1) | 1 |
| 2021 | Latent Compositional Representations Improve Systematic Generalization in Grounded Question AnsweringabstractAbstract Answering questions that involve multi-step reasoning requires decomposing them and using the answers of intermediate steps to reach the final answer. However, state-of-the-art models in grounded question answering often do not explicitly perform decomposition, leading to difficulties in generalization to out-of-distribution examples. In this work, we propose a model that computes a representation and denotation for all question spans in a bottom-up, compositional manner using a CKY-style parser. Our model induces latent trees, driven by end-to-end (the answer) supervision only. We show that this inductive bias towards tree structures dramatically improves systematic generalization to out-of- distribution examples, compared to strong baselines on an arithmetic expressions benchmark as well as on C losure, a dataset that focuses on systematic generalization for grounded question answering. On this challenging dataset, our model reaches an accuracy of 96.1%, significantly higher than prior models that almost perfectly solve the task on a random, in-distribution split. Ben Bogin, Sanjay Subramanian, Matt Gardner 0001, Jonathan Berant |
Trans. Assoc. Comput. Linguistics | 2 |
| 2020 | Obtaining Faithful Interpretations from Compositional Neural NetworksabstractNeural module networks (NMNs) are a popular approach for modeling compositionality: they achieve high accuracy when applied to problems in language and vision, while reflecting the compositional structure of the problem in the network architecture.However, prior work implicitly assumed that the structure of the network modules, describing the abstract reasoning process, provides a faithful explanation of the model's reasoning; that is, that all modules perform their intended behaviour.In this work, we propose and conduct a systematic evaluation of the intermediate outputs of NMNs on NLVR2 and DROP, two datasets which require composing multiple reasoning steps.We find that the intermediate outputs differ from the expected output, illustrating that the network structure does not provide a faithful explanation of model behaviour.To remedy that, we train the model with auxiliary supervision and propose particular choices for module architecture that yield much better faithfulness, at a minimal cost to accuracy. Sanjay Subramanian, Ben Bogin, Nitish Gupta, Tomer Wolfson, Sameer Singh 0001, Jonathan Berant, Matt Gardner 0001 |
ACL | 1 |
| 2019 | An Improved Neural Baseline for Temporal Relation ExtractionabstractQiang Ning, Sanjay Subramanian, Dan Roth. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Qiang Ning, Sanjay Subramanian, Dan Roth 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Correlation Clustering with Same-Cluster Queries Bounded by Optimal CostabstractSeveral clustering frameworks with interactive (semi-supervised) queries have been studied in the past. Recently, clustering with same-cluster queries has become popular. An algorithm in this setting has access to an oracle with full knowledge of an optimal clustering, and the algorithm can ask the oracle queries of the form, "Does the optimal clustering put vertices u and v in the same cluster?" Due to its simplicity, this querying model can easily be implemented in real crowd-sourcing platforms and has attracted a lot of recent work. In this paper, we study the popular correlation clustering problem (Bansal et al., 2002) under the same-cluster querying framework. Given a complete graph G=(V,E) with positive and negative edge labels, correlation clustering objective aims to compute a graph clustering that minimizes the total number of disagreements, that is the negative intra-cluster edges and positive inter-cluster edges. In a recent work, Ailon et al. (2018b) provided an approximation algorithm for correlation clustering that approximates the correlation clustering objective within (1+epsilon) with O((k^{14} log{n} log{k})/epsilon^6) queries when the number of clusters, k, is fixed. For many applications, k is not fixed and can grow with |V|. Moreover, the dependency of k^14 on query complexity renders the algorithm impractical even for datasets with small values of k. In this paper, we take a different approach. Let C_{OPT} be the number of disagreements made by the optimal clustering. We present algorithms for correlation clustering whose error and query bounds are parameterized by C_{OPT} rather than by the number of clusters. Indeed, a good clustering must have small C_{OPT}. Specifically, we present an efficient algorithm that recovers an exact optimal clustering using at most 2C_{OPT} queries and an efficient algorithm that outputs a 2-approximation using at most C_{OPT} queries. In addition, we show under a plausible complexity assumption, there does not exist any polynomial time algorithm that has an approximation ratio better than 1+alpha for an absolute constant alpha > 0 with o(C_{OPT}) queries. Therefore, our first algorithm achieves the optimal query bound within a factor of 2. We extensively evaluate our methods on several synthetic and real-world datasets using real crowd-sourced oracles. Moreover, we compare our approach against known correlation clustering algorithms that do not perform querying. In all cases, our algorithms exhibit superior performance. Barna Saha, Sanjay Subramanian |
ESA | 2 |