Shivang Chopra

dblp:262/6038 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0002-3567-852XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Trustworthy machine learning · 50% Transfer learning and domain adaptation · 30% Vision and language · 20%
Databases, data mining, and information retrieval
1 paper
Web and social media mining · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation › fine-tuning
robust fine-tuning
1.722025
Directional Gradient Projection for Robust Fine-Tuning of Foundation Models · ICLR 2025
FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering · CVPR 2025
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift
1.722025
Directional Gradient Projection for Robust Fine-Tuning of Foundation Models · ICLR 2025
FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering · CVPR 2025
Computer vision › Vision and language
visual question answering
1.122025
FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering · CVPR 2025
Directional Gradient Projection for Robust Fine-Tuning of Foundation Models · ICLR 2025
Machine learning › Trustworthy machine learning
debiasing
0.412020
Hindi-English Hate Speech Detection: Author Profiling, Debiasing, and Practical Perspectives · AAAI 2020
Machine learning › Trustworthy machine learning
fairness
0.412020
Hindi-English Hate Speech Detection: Author Profiling, Debiasing, and Practical Perspectives · AAAI 2020
Web and social media mining › hate speech
hate speech detection
0.412020
Hindi-English Hate Speech Detection: Author Profiling, Debiasing, and Practical Perspectives · AAAI 2020
Machine learning › Trustworthy machine learning
out-of-distribution generalization
0.312025
FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering · CVPR 2025

Methods — techniques the papers use, named apart from their topics

regularization · 0.9multi-objective optimization · 0.9mahalanobis distance · 0.9gradient projection · 0.9embedding analysis · 0.9profanity modeling · 0.9graph embedding · 0.9expert-in-the-loop debiasing · 0.9
YearPublicationVenuePosition
2025 FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering
abstract
Visual question answering (VQA) systems face significant challenges when adapting to real-world data shifts, especially in multi-modal contexts. While robust fine-tuning strategies are essential for maintaining performance across in-distribution (ID) and out-of-distribution (OOD) scenarios, current evaluation settings are primarily unimodal or particular to some types of OOD, offering limited insight into the complexities of multi-modal contexts. In this work, we propose a new benchmark FRAMES-VQA (Fine-Tuning Robustness Across Multi-Modal Shifts in VQA) for evaluating robust fine-tuning for VQA tasks. We utilize ten existing VQA benchmarks, including VQAv2, IV-VQA, VQACP, OK-VQA and others, and categorize them into ID, near and far OOD datasets covering uni-modal, multi-modal and adversarial distribution shifts. We first conduct a comprehensive comparison of existing robust fine-tuning methods. We then quantify the distribution shifts by calculating the Mahalanobis distance using uni-modal and multimodal embeddings extracted from various models. Further, we perform an extensive analysis to explore the interactions between uni- and multi-modal shifts as well as modality importance for ID and OOD samples. These analyses offer valuable guidance on developing more robust fine-tuning methods to handle multi-modal distribution shifts. The code is available at https://github.com/chengyuehuang511/FRAMES-VQA.
Chengyue Huang, Brisa Maneechotesuwan, Shivang Chopra, Zsolt Kira
CVPR3
2025 Directional Gradient Projection for Robust Fine-Tuning of Foundation Models
abstract
Robust fine-tuning aims to adapt large foundation models to downstream tasks while preserving their robustness to distribution shifts. Existing methods primarily focus on constraining and projecting current model towards the pre-trained initialization based on the magnitudes between fine-tuned and pre-trained weights, which often require extensive hyper-parameter tuning and can sometimes result in underfitting. In this work, we propose $\textbf{Di}$rectional $\textbf{Gra}$dient $\textbf{P}$rojection (DiGraP), a novel layer-wise trainable method that incorporates directional information from gradients to bridge regularization and multi-objective optimization. Besides demonstrating our method on image classification, as another contribution we generalize this area to the multi-modal evaluation settings for robust fine-tuning. Specifically, we first bridge the uni-modal and multi-modal gap by performing analysis on Image Classification reformulated Visual Question Answering (VQA) benchmarks and further categorize ten out-of-distribution (OOD) VQA datasets by distribution shift types and degree (i.e. near versus far OOD). Experimental results show that DiGraP consistently outperforms existing baselines across Image Classfication and VQA tasks with discriminative and generative backbones, improving both in-distribution (ID) generalization and OOD robustness.
Chengyue Huang, Junjiao Tian, Brisa Maneechotesuwan, Shivang Chopra, Zsolt Kira
ICLR4
2025 Refining Text-to-Image Generation: Towards Accurate Training-Free Glyph-Enhanced Image Generation
abstract
Over the past few years, Text-to-Image (T2I) generation approaches based on diffusion models have gained signifi-cant attention. However, vanilla diffusion models often suffer from spelling inaccuracies in the text displayed within the generated images. The capability to generate visual text is crucial, offering both academic interest and a wide range of practical applications. To produce accurate visual text images, state-of-the-art techniques adopt a glyph-controlled image generation approach, consisting of a text layout generator followed by an image generator that is conditioned on the generated text layout. Nevertheless, our study reveals that these models still face three primary challenges, prompting us to develop a testbed to facilitate future research. We introduce a benchmark, LenCom-Eval, specifically designed for testing models' capability in generating images with Lengthy and Complex visual text. Subsequently, we introduce a training-free framework to enhance the two-stage generation approaches. We examine the effectiveness of our approach on both LenCom-Eval and MARIO-Eval bench-marks and demonstrate notable improvements across a range of evaluation metrics, including CLIPScore, OCR precision, recall, F1 score, accuracy, and edit distance scores. For instance, our proposed framework improves the backbone model, TextDiffuser, by more than 23% and 13.5% in terms of OCR word F1 on LenCom-Eval and MARIO-Eval, re-spectively. Our work makes a unique contribution to the field by focusing on generating images with long and rare text sequences, a niche previously unexplored by existing literature11The code and the datasets are available at https://github.com/SanyamLakhanpal/SAOcr/tree/main.
Sanyam Lakhanpal, Shivang Chopra, Vinija Jain, Aman Chadha
WACV2
2021 SketchFormer: transformer-based approach for sketch recognition using vector images
Anil Singh Parihar, Gaurav Jain, Shivang Chopra, Suransh Chopra
Multim. Tools Appl.3
2020 Hindi-English Hate Speech Detection: Author Profiling, Debiasing, and Practical Perspectives
abstract
Code-switching in linguistically diverse, low resource languages is often semantically complex and lacks sophisticated methodologies that can be applied to real-world data for precisely detecting hate speech. In an attempt to bridge this gap, we introduce a three-tier pipeline that employs profanity modeling, deep graph embeddings, and author profiling to retrieve instances of hate speech in Hindi-English code-switched language (Hinglish) on social media platforms like Twitter. Through extensive comparison against several baselines on two real-world datasets, we demonstrate how targeted hate embeddings combined with social network-based features outperform state of the art, both quantitatively and qualitatively. Additionally, we present an expert-in-the-loop algorithm for bias elimination in the proposed model pipeline and study the prevalence and performance impact of the debiasing. Finally, we discuss the computational, practical, ethical, and reproducibility aspects of the deployment of our pipeline across the Web.
Shivang Chopra, Ramit Sawhney, Puneet Mathur, Rajiv Ratn Shah
AAAI1
2020 TransSketchNet: Attention-Based Sketch Recognition Using Transformers
Gaurav Jain, Shivang Chopra, Suransh Chopra, Anil Singh Parihar
ECAI2
2020 Utilizing Temporal Psycholinguistic Cues for Suicidal Intent Estimation
Puneet Mathur, Ramit Sawhney, Shivang Chopra, Maitree Leekha, Rajiv Ratn Shah
ECIR (2)3