Agneet Chatterjee

dblp:218/8410 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0002-0961-9569ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models
abstract
Text-to-Image (T2I) models have recently achieved remarkable success in generating images from textual descriptions.However, challenges still persist in accurately rendering complex scenes where actions and interactions form the primary semantic focus.Our key observation in this work is that T2I models frequently struggle to capture nuanced and often implicit attributes inherent in action depiction, leading to generating images that lack key contextual details.To enable systematic evaluation, we introduce AcT2I, a benchmark designed to evaluate the performance of T2I models in generating images from action-centric prompts.We experimentally validate that leading T2I models do not fare well on AcT2I.We further hypothesize that this shortcoming arises from the incomplete representation of the inherent attributes and contextual dependencies in the training corpora of existing T2I models.We build upon this by developing a trainingfree, knowledge distillation technique utilizing Large Language Models to address this limitation.Specifically, we enhance prompts by incorporating dense information across three dimensions, observing that injecting prompts with temporal details significantly improves image generation accuracy, with our best model achieving an increase of 72%.Our findings highlight the limitations of current T2I methods in generating images that require complex reasoning and demonstrate that integrating linguistic knowledge in a systematic way can notably advance the generation of nuanced and contextually accurate images.
Vatsal Malaviya, Agneet Chatterjee, Maitreya Patel, Yezhou Yang, Chitta Baral
EMNLP2
2025 Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation
abstract
Recent advances in video generation have enabled high-fidelity video synthesis from user provided prompts. However, existing models and benchmarks fail to capture the complexity and requirements of professional video generation. Towards that goal, we introduce Stable Cinemetrics, a structured evaluation framework that formalizes filmmaking controls into four disentangled, hierarchical taxonomies: Setup, Event, Lighting, and Camera. Together, these taxonomies define 76 fine-grained control nodes grounded in industry practices. Using these taxonomies, we construct a benchmark of prompts aligned with professional use cases and develop an automated pipeline for prompt categorization and question generation, enabling independent evaluation of each control dimension. We conduct a large-scale human study spanning 10+ models and 20K videos, annotated by a pool of 80+ film professionals. Our analysis, both coarse and fine-grained reveal that even the strongest current models exhibit significant gaps, particularly in Events and Camera-related controls. To enable scalable evaluation, we train an automatic evaluator, a vision-language model aligned with expert annotations that outperforms existing zero-shot baselines. SCINE is the first approach to situate professional video generation within the landscape of video generative models, introducing taxonomies centered around cinematic controls and supporting them with structured evaluation pipelines and detailed analyses to guide future research.
Agneet Chatterjee, Rahim Entezari, Maksym Zhuravinskyi, Maksim Lapin, Reshinth Adithyan, Amit Raj, Chitta Baral, Yezhou Yang, Varun Jampani
NeurIPS1
2024 On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
abstract
Recent advances in monocular depth estimation have been made by incorporating natural language as additional guidance. Although yielding impressive results, the impact of the language prior, particularly in terms of generalization and robustness, remains unexplored. In this paper, we address this gap by quantifying the impact of this prior and introduce methods to benchmark its effectiveness across various settings. We generate “low-level” sentences that convey object-centric, three-dimensional spatial relationships, incorporate them as additional language priors and evaluate their downstream impact on depth estimation. Our key finding is that current language-guided depth estimators perform optimally only with scene-level descriptions and counter-intuitively fare worse with low level descriptions. Despite leveraging additional data, these methods are not robust to directed adversarial attacks and decline in performance with an increase in distribution shift. Finally, to provide a foundation for future research, we identify points of failures and offer insights to better understand these shortcomings. With an increasing number of methods using language for depth estimation, our findings highlight the opportunities and pitfalls that require careful consideration for effective deployment in real-world settings.11Code/Data: https://github.com/agneet42/lang_depth
Agneet Chatterjee, Tejas Gokhale, Chitta Baral, Yezhou Yang
CVPR1
2024 REVISION: Rendering Tools Enable Spatial Fidelity in Vision-Language Models
Agneet Chatterjee, Yiran Luo 0001, Tejas Gokhale, Yezhou Yang, Chitta Baral
ECCV (30)1
2024 Getting it Right: Improving Spatial Consistency in Text-to-Image Models
Agneet Chatterjee, Gabriela Ben Melech Stan, Estelle Aflalo, Sayak Paul, Dhruba Ghosh, Tejas Gokhale, Ludwig Schmidt, Hannaneh Hajishirzi, Vasudev Lal, Chitta Baral, Yezhou Yang
ECCV (22)1
2023 A novel meta-heuristic approach for influence maximization in social networks
abstract
Abstract Influence maximization in a social network focuses on the task of extracting a small set of nodes from a network which can maximize the propagation in a cascade model. Though greedy methods produce good solutions to the aforementioned problem, their high computational complexity is a major drawback. Centrality‐based heuristic methods often fail to overcome local optima, thereby producing sub‐optimal results. To this end, in this article, a framework has been presented which involves community detection in a social network and the utilization of the Shuffled Frog Leaping algorithm, in maximizing the two‐hop spread of influence under the independent cascade model. Local search strategies like the Late acceptance based hill climbing have been employed to improve the solution further. Experiments performed on three real‐world datasets have shown that our method performs markedly well with respect to the comparing algorithms.
Bitanu Chatterjee, Trinav Bhattacharyya, Kushal Kanti Ghosh, Agneet Chatterjee, Ram Sarkar
Expert Syst. J. Knowl. Eng.4
2021 A two-phase gradient based feature embedding approach
Agneet Chatterjee, Soulib Ghosh, Anuran Chakraborty, Sudipta Kumar Ghosal, Ram Sarkar
J. Inf. Secur. Appl.1
2021 Image steganography based on Kirsch edge detection
Sudipta Kumar Ghosal, Agneet Chatterjee, Ram Sarkar
Multim. Syst.2
2021 Application of daisy descriptor for language identification in the wild
Neelotpal Chakraborty, Agneet Chatterjee, Pawan Kumar Singh 0001, Ayatullah Faruk Mollah, Ram Sarkar
Multim. Tools Appl.2
2021 CTRL -CapTuRedLight: a novel feature descriptor for online Assamese numeral recognition
Soulib Ghosh, Agneet Chatterjee, Shibaprasad Sen, Neeraj Kumar 0001, Ram Sarkar
Multim. Tools Appl.2
2021 An ensemble approach to outlier detection using some conventional clustering algorithms
Agneet Chatterjee, Soulib Ghosh, Neeraj Kumar 0001, Ram Sarkar
Multim. Tools Appl.2
2021 Language-invariant novel feature descriptors for handwritten numeral recognition
Soulib Ghosh, Agneet Chatterjee, Pawan Kumar Singh 0001, Showmik Bhowmik, Ram Sarkar
Vis. Comput.2
2020 LSB based steganography with OCR: an intelligent amalgamation
Agneet Chatterjee, Sudipta Kumar Ghosal, Ram Sarkar
Multim. Tools Appl.1
2020 Offline music symbol recognition using Daisy feature and quantum Grey wolf optimization based feature selection
Samir Malakar, Manosij Ghosh, Agneet Chatterjee, Showmik Bhowmik, Ram Sarkar
Multim. Tools Appl.3
2020 Extended exploiting modification direction based steganography using hashed-weightage Array
Shaswata Saha, Anuran Chakraborty, Agneet Chatterjee, Souvik Dhargupta, Sudipta Kumar Ghosal, Ram Sarkar
Multim. Tools Appl.3
2019 Filter Method Ensemble with Neural Networks
Anuran Chakraborty, Rajonya De, Agneet Chatterjee, Friedhelm Schwenker, Ram Sarkar
ICANN (2)3