Anil Batra

dblp:226/5040 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Information extraction and text analysis · 31% Vision and language · 29% Generative modeling · 23%
Computer graphics and multimedia
1 paper
Image and video processing · 50% Multimedia analysis and retrieval · 50%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › semantic role labeling
implicit argument prediction
0.912025
Predicting Implicit Arguments in Procedural Video Instructions · ACL (1) 2025
Natural language and speech › Information extraction and text analysis
semantic role labeling
0.912025
Predicting Implicit Arguments in Procedural Video Instructions · ACL (1) 2025
Machine learning › Generative modeling
diffusion model
0.712023
Image generation with shortest path diffusion · ICML 2023
Machine learning › Generative modeling
image generation
0.712023
Image generation with shortest path diffusion · ICML 2023
Robotics › Motion planning and robot control › robot learning › object learning
orientation learning
0.412019
Improved Road Connectivity by Joint Learning of Orientation and Segmentation · CVPR 2019
Computer vision › Segmentation and scene understanding
semantic segmentation
0.412019
Improved Road Connectivity by Joint Learning of Orientation and Segmentation · CVPR 2019
Image and video processing
image segmentation
0.412019
Improved Road Connectivity by Joint Learning of Orientation and Segmentation · CVPR 2019
Multimedia analysis and retrieval › image analysis › aerial image analysis
road network extraction
0.412019
Improved Road Connectivity by Joint Learning of Orientation and Segmentation · CVPR 2019
Computer vision › Video understanding and tracking › activity recognition › procedural activity understanding
procedural video understanding
0.212024
Efficient Pre-training for Localized Instruction Generation of Procedural Videos · ECCV (39) 2024

Methods — techniques the papers use, named apart from their topics

multimodal large language model · 0.9entity tracking · 0.9pre-training · 0.8multi-branch convolutional module · 0.8connectivity refinement · 0.8information geometry · 0.7fisher metric · 0.7
YearPublicationVenuePosition
2025 Predicting Implicit Arguments in Procedural Video Instructions
abstract
Procedural texts help AI enhance reasoning about context and action sequences. Transforming these into Semantic Role Labeling (SRL) improves understanding of individual steps by identifying predicate-argument structure like verb,what,where/with. Procedural instructions are highly elliptic, for instance, (i) add cucumber to the bowl and (ii) add sliced tomatoes, the second step's where argument is inferred from the context, referring to where the cucumber was placed. Prior SRL benchmarks often miss implicit arguments, leading to incomplete understanding. To address this, we introduce Implicit-VidSRL, a dataset that necessitates inferring implicit and explicit arguments from contextual information in multimodal cooking procedures. Our proposed dataset benchmarks multimodal models' contextual reasoning, requiring entity tracking through visual changes in recipes. We study recent multimodal LLMs and reveal that they struggle to predict implicit arguments of what and where/with from multi-modal procedural data given the verb. Lastly, we propose iSRL-Qwen2-VL, which achieves a 17% relative improvement in F1-score for what-implicit and a 14.7% for where/with-implicit semantic roles over GPT-4o.
Anil Batra, Laura Sevilla-Lara, Marcus Rohrbach, Frank Keller
ACL (1)1
2025 CAST: Cross-modal Alignment Similarity Test for Vision Language Models
abstract
Vision Language Models (VLMs) are typically evaluated with Visual Question Answering (VQA) tasks which assess a model’s understanding of scenes. Good VQA performance is taken as evidence that the model will perform well on a broader range of tasks that require both visual and language inputs. However, scene-aware VQA does not fully capture input biases or assess hallucinations caused by a misalignment between modalities. To address this, we propose a Cross-modal Alignment Similarity Test (CAST) to probe VLMs for self-consistency across modalities. This test involves asking the models to identify similarities between two scenes through text-only, image-only, or both and then assess the truthfulness of the similarities they generate. Since there is no ground-truth to compare against, this evaluation does not focus on objective accuracy but rather on whether VLMs are internally consistent in their outputs. We argue that while not all self-consistent models are capable or accurate, all capable VLMs must be self-consistent.
Gautier Dagan, Olga Loginova, Anil Batra
COLING3
2024 Efficient Pre-training for Localized Instruction Generation of Procedural Videos
Anil Batra, Davide Moltisanti, Laura Sevilla-Lara, Marcus Rohrbach, Frank Keller
ECCV (39)1
2023 Image generation with shortest path diffusion
abstract
The field of image generation has made significant progress thanks to the introduction of Diffusion Models, which learn to progressively reverse a given image corruption. Recently, a few studies introduced alternative ways of corrupting images in Diffusion Models, with an emphasis on blurring. However, these studies are purely empirical and it remains unclear what is the optimal procedure for corrupting an image. In this work, we hypothesize that the optimal procedure minimizes the length of the path taken when corrupting an image towards a given final state. We propose the Fisher metric for the path length, measured in the space of probability distributions. We compute the shortest path according to this metric, and we show that it corresponds to a combination of image sharpening, rather than blurring, and noise deblurring. While the corruption was chosen arbitrarily in previous work, our Shortest Path Diffusion (SPD) determines uniquely the entire spatiotemporal structure of the corruption. We show that SPD improves on strong baselines without any hyperparameter tuning, and outperforms all previous Diffusion Models based on image blurring. Furthermore, any small deviation from the shortest path leads to worse performance, suggesting that SPD provides the optimal procedure to corrupt images. Our work sheds new light on observations made in recent works and provides a new approach to improve diffusion models on images and other types of data.
Ayan Das 0005, Stathi Fotiadis, Anil Batra, Farhang Nabiei, Fengting Liao, Sattar Vakili, Da-Shan Shiu, Alberto Bernacchia
ICML3
2022 A Closer Look at Temporal Ordering in the Segmentation of Instructional Videos
Anil Batra, Shreyank N. Gowda, Frank Keller, Laura Sevilla-Lara
BMVC1
2019 Improved Road Connectivity by Joint Learning of Orientation and Segmentation
abstract
Road network extraction from satellite images often produce fragmented road segments leading to road maps unfit for real applications. Pixel-wise classification fails to predict topologically correct and connected road masks due to the absence of connectivity supervision and difficulty in enforcing topological constraints. In this paper, we propose a connectivity task called Orientation Learning, motivated by the human behavior of annotating roads by tracing it at a specific orientation. We also develop a stacked multi-branch convolutional module to effectively utilize the mutual information between orientation learning and segmentation tasks. These contributions ensure that the model predicts topologically correct and connected road masks. We also propose Connectivity Refinement approach to further enhance the estimated road networks. The refinement model is pre-trained to connect and refine the corrupted ground-truth masks and later fine-tuned to enhance the predicted road masks. We demonstrate the advantages of our approach on two diverse road extraction datasets SpaceNet and DeepGlobe. Our approach improves over the state-of-the-art techniques by 9% and 7.5% in road topology metric on SpaceNet and DeepGlobe, respectively.
Anil Batra, Suriya Singh, Guan Pang, Saikat Basu, C. V. Jawahar, Manohar Paluri
CVPR1
2018 Self-Supervised Feature Learning for Semantic Segmentation of Overhead Imagery
Suriya Singh, Anil Batra, Guan Pang, Lorenzo Torresani, Saikat Basu, Manohar Paluri, C. V. Jawahar
BMVC2