Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Satya Narayan Shukla

dblp:161/3356 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
7since 2021 · last 2025
0009-0009-3982-4414ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Vision and language · 42% Time series and sequential data · 30% Deep learning architectures and training · 9%
Network and information security
1 paper
Security and privacy of machine learning · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
vision-language model
1.522024
uCAP: An Unsupervised Prompting Method for Vision-Language Models · ECCV (74) 2024
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs · CVPR 2024
Machine learning › Time series and sequential data › time series modeling
irregularly sampled time series
1.532022
Heteroscedastic Temporal Variational Autoencoder For Irregularly Sampled Time Series · ICLR 2022
Multi-Time Attention Networks for Irregularly Sampled Time Series · ICLR 2021
Interpolation-Prediction Networks for Irregularly Sampled Time Series · ICLR (Poster) 2019
Machine learning › Time series and sequential data
time series modeling
1.532022
Heteroscedastic Temporal Variational Autoencoder For Irregularly Sampled Time Series · ICLR 2022
Multi-Time Attention Networks for Irregularly Sampled Time Series · ICLR 2021
Interpolation-Prediction Networks for Irregularly Sampled Time Series · ICLR (Poster) 2019
Computer vision › Vision and language
cross-modal alignment
0.912025
CompCap: Improving Multimodal Large Language Models with Composite Captions · ICCV 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
CompCap: Improving Multimodal Large Language Models with Composite Captions · ICCV 2025
Natural language and speech › Question answering and dialogue systems
machine reading comprehension
0.812024
The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants · ACL (1) 2024
Computer vision › Vision and language
visual question answering
0.812024
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs · CVPR 2024
Machine learning › Generative modeling
variational autoencoder
0.612022
Heteroscedastic Temporal Variational Autoencoder For Irregularly Sampled Time Series · ICLR 2022
Machine learning › Deep learning architectures and training
attention mechanism
0.512021
Multi-Time Attention Networks for Irregularly Sampled Time Series · ICLR 2021
Security and privacy of machine learning
adversarial attack
0.512021
Simple and Efficient Hard Label Black-box Adversarial Attacks in Low Query Budget Regimes · KDD 2021
Security and privacy of machine learning › adversarial attack
black-box attack
0.512021
Simple and Efficient Hard Label Black-box Adversarial Attacks in Low Query Budget Regimes · KDD 2021
Security and privacy of machine learning › adversarial attack › black-box attack
hard-label attack
0.512021
Simple and Efficient Hard Label Black-box Adversarial Attacks in Low Query Budget Regimes · KDD 2021
Security and privacy of machine learning › adversarial attack
query-efficient attack
0.512021
Simple and Efficient Hard Label Black-box Adversarial Attacks in Low Query Budget Regimes · KDD 2021
Machine learning › Generative modeling
synthetic data generation
0.312025
CompCap: Improving Multimodal Large Language Models with Composite Captions · ICCV 2025
Computer vision › Image recognition and object detection
object localization
0.212024
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs · CVPR 2024
Information retrieval
evaluation
0.212024
The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants · ACL (1) 2024
Information retrieval › evaluation › benchmark
multilingual benchmark
0.212024
The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants · ACL (1) 2024

Methods — techniques the papers use, named apart from their topics

dataset construction · 1.5supervised fine-tuning · 0.9large language model · 0.9pseudo data generation · 0.8prompting · 0.8instruction fine-tuning · 0.8coordinate-based supervision · 0.8variational inference · 0.6heteroscedastic modeling · 0.6low-dimensional subspace search · 0.5bayesian optimization · 0.5attention · 0.5
YearPublicationVenuePosition
2025 CompCap: Improving Multimodal Large Language Models with Composite Captions
abstract
How well can Multimodal Large Language Models (MLLMs) understand composite images? Composite images (CIs) are synthetic visuals created by merging multiple visual elements, such as charts, posters, or screenshots, rather than being captured directly by a camera. While CIs are prevalent in real-world applications, recent MLLM developments have primarily focused on interpreting natural images (NIs). Our research reveals that current MLLMs face significant challenges in accurately understanding CIs, often struggling to extract information or perform complex reasoning based on these images. We find that existing training data for CIs are mostly formatted for question-answer tasks (e.g., in datasets like ChartQA and ScienceQA), while high-quality image-caption datasets, critical for robust vision-language alignment, are only available for NIs. To bridge this gap, we introduce Composite Captions (CompCap), a flexible framework that leverages Large Language Models (LLMs) and automation tools to synthesize CIs with accurate and detailed captions. Using CompCap, we curate CompCap-118K, a dataset containing 118K image-caption pairs across six CI types. We validate the effectiveness of CompCap-118K by supervised fine-tuning MLLMs of three sizes: xGen-MM-inst.-4B and LLaVA-NeXT-Vicuna-7B/13B. Empirical results show that CompCap-118K significantly enhances MLLMs' understanding of CIs, yielding average gains of 1.7%, 2.0%, and 2.9% across eleven benchmarks, respectively.
Satya Narayan Shukla, Mahmoud Azab, Aashu Singh, Qifan Wang 0001, Shengyun Peng, Hanchao Yu, Shen Yan 0007, Xuewen Zhang, Baosheng He
ICCV2
2024 The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants
abstract
Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, Madian Khabsa. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal 0001, Abhinandan Krishnan, Luke Zettlemoyer, Madian Khabsa
ACL (1)5
2024 Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs
abstract
Integration of Large Language Models (LLMs) into visual domain tasks, resulting in visual-LLMs (V-LLMs), has enabled exceptional performance in vision-language tasks, particularly for visual question answering (VQA). However, existing V-LLMs (e.g. BLIP-2, LLaVA) demonstrate weak spatial reasoning and localization awareness. Despite generating highly descriptive and elaborate textual answers, these models fail at simple tasks like distinguishing a left vs right location. In this work, we explore how image-space coordinate based instruction fine-tuning objectives could inject spatial awareness into V-LLMs. We discover optimal coordinate representations, data-efficient instruction fine-tuning objectives, and pseudo-data generation strategies that lead to improved spatial awareness in V-LLMs. Additionally, our resulting model improves VQA across image and video domains, reduces undesired hallucination, and generates better contextual object descriptions. Experiments across 5 vision-language tasks involving 14 different datasets establish the clear performance improvements achieved by our proposed framework.
Kanchana Ranasinghe, Satya Narayan Shukla, Omid Poursaeed, Michael S. Ryoo, Tsung-Yu Lin
CVPR2
2024 uCAP: An Unsupervised Prompting Method for Vision-Language Models
A. Tuan Nguyen, Kai Sheng Tai, Bor-Chun Chen, Satya Narayan Shukla, Hanchao Yu, Philip Torr 0001, Tai-Peng Tian, Ser-Nam Lim
ECCV (74)4
2022 Heteroscedastic Temporal Variational Autoencoder For Irregularly Sampled Time Series
Satya Narayan Shukla, Benjamin M. Marlin
ICLR1
2021 Multi-Time Attention Networks for Irregularly Sampled Time Series
Satya Narayan Shukla, Benjamin M. Marlin
ICLR1
2021 Simple and Efficient Hard Label Black-box Adversarial Attacks in Low Query Budget Regimes
abstract
We focus on the problem of black-box adversarial attacks, where the aim is to generate adversarial examples for deep learning models solely based on information limited to output label (hard-label) to a queried data input. We propose a simple and efficient Bayesian Optimization (BO) based approach for developing black-box adversarial attacks. Issues with BO's performance in high dimensions are avoided by searching for adversarial examples in a structured low-dimensional subspace. We demonstrate the efficacy of our proposed attack method by evaluating both ℓ∞ and ℓ2 norm constrained untargeted and targeted hard label black-box attacks on three standard datasets - MNIST, CIFAR-10, and ImageNet. Our proposed approach consistently achieves 2x to 10x higher attack success rate while requiring 10x to 20x fewer queries compared to the current state-of-the-art black-box adversarial attacks.
Satya Narayan Shukla, Anit Kumar Sahu, Devin Willmott, J. Zico Kolter
KDD1
2019 Interpolation-Prediction Networks for Irregularly Sampled Time Series
Satya Narayan Shukla, Benjamin M. Marlin
ICLR (Poster)1