EDBT 2026 Demo / reviewers in the wild / expert
Satya Narayan Shukla
dblp:161/3356
· DBLP profile ↗
8ranked-venue papers
4as first author
7since 2021 · last 2025
0009-0009-3982-4414ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Vision and language · 42% Time series and sequential data · 30% Deep learning architectures and training · 9% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
vision-language model |
1.5 | 2 | 2024 | uCAP: An Unsupervised Prompting Method for Vision-Language Models · ECCV (74) 2024 Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs · CVPR 2024 |
Machine learning › Time series and sequential data › time series modeling
irregularly sampled time series |
1.5 | 3 | 2022 | Heteroscedastic Temporal Variational Autoencoder For Irregularly Sampled Time Series · ICLR 2022 Multi-Time Attention Networks for Irregularly Sampled Time Series · ICLR 2021 Interpolation-Prediction Networks for Irregularly Sampled Time Series · ICLR (Poster) 2019 |
Machine learning › Time series and sequential data
time series modeling |
1.5 | 3 | 2022 | Heteroscedastic Temporal Variational Autoencoder For Irregularly Sampled Time Series · ICLR 2022 Multi-Time Attention Networks for Irregularly Sampled Time Series · ICLR 2021 Interpolation-Prediction Networks for Irregularly Sampled Time Series · ICLR (Poster) 2019 |
Computer vision › Vision and language
cross-modal alignment |
0.9 | 1 | 2025 | CompCap: Improving Multimodal Large Language Models with Composite Captions · ICCV 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | CompCap: Improving Multimodal Large Language Models with Composite Captions · ICCV 2025 |
Natural language and speech › Question answering and dialogue systems
machine reading comprehension |
0.8 | 1 | 2024 | The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants · ACL (1) 2024 |
Computer vision › Vision and language
visual question answering |
0.8 | 1 | 2024 | Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs · CVPR 2024 |
Machine learning › Generative modeling
variational autoencoder |
0.6 | 1 | 2022 | Heteroscedastic Temporal Variational Autoencoder For Irregularly Sampled Time Series · ICLR 2022 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.5 | 1 | 2021 | Multi-Time Attention Networks for Irregularly Sampled Time Series · ICLR 2021 |
Security and privacy of machine learning
adversarial attack |
0.5 | 1 | 2021 | Simple and Efficient Hard Label Black-box Adversarial Attacks in Low Query Budget Regimes · KDD 2021 |
Security and privacy of machine learning › adversarial attack
black-box attack |
0.5 | 1 | 2021 | Simple and Efficient Hard Label Black-box Adversarial Attacks in Low Query Budget Regimes · KDD 2021 |
Security and privacy of machine learning › adversarial attack › black-box attack
hard-label attack |
0.5 | 1 | 2021 | Simple and Efficient Hard Label Black-box Adversarial Attacks in Low Query Budget Regimes · KDD 2021 |
Security and privacy of machine learning › adversarial attack
query-efficient attack |
0.5 | 1 | 2021 | Simple and Efficient Hard Label Black-box Adversarial Attacks in Low Query Budget Regimes · KDD 2021 |
Machine learning › Generative modeling
synthetic data generation |
0.3 | 1 | 2025 | CompCap: Improving Multimodal Large Language Models with Composite Captions · ICCV 2025 |
Computer vision › Image recognition and object detection
object localization |
0.2 | 1 | 2024 | Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs · CVPR 2024 |
Information retrieval
evaluation |
0.2 | 1 | 2024 | The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants · ACL (1) 2024 |
Information retrieval › evaluation › benchmark
multilingual benchmark |
0.2 | 1 | 2024 | The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants · ACL (1) 2024 |
Methods — techniques the papers use, named apart from their topics
dataset construction · 1.5supervised fine-tuning · 0.9large language model · 0.9pseudo data generation · 0.8prompting · 0.8instruction fine-tuning · 0.8coordinate-based supervision · 0.8variational inference · 0.6heteroscedastic modeling · 0.6low-dimensional subspace search · 0.5bayesian optimization · 0.5attention · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CompCap: Improving Multimodal Large Language Models with Composite CaptionsabstractHow well can Multimodal Large Language Models (MLLMs) understand composite images? Composite images (CIs) are synthetic visuals created by merging multiple visual elements, such as charts, posters, or screenshots, rather than being captured directly by a camera. While CIs are prevalent in real-world applications, recent MLLM developments have primarily focused on interpreting natural images (NIs). Our research reveals that current MLLMs face significant challenges in accurately understanding CIs, often struggling to extract information or perform complex reasoning based on these images. We find that existing training data for CIs are mostly formatted for question-answer tasks (e.g., in datasets like ChartQA and ScienceQA), while high-quality image-caption datasets, critical for robust vision-language alignment, are only available for NIs. To bridge this gap, we introduce Composite Captions (CompCap), a flexible framework that leverages Large Language Models (LLMs) and automation tools to synthesize CIs with accurate and detailed captions. Using CompCap, we curate CompCap-118K, a dataset containing 118K image-caption pairs across six CI types. We validate the effectiveness of CompCap-118K by supervised fine-tuning MLLMs of three sizes: xGen-MM-inst.-4B and LLaVA-NeXT-Vicuna-7B/13B. Empirical results show that CompCap-118K significantly enhances MLLMs' understanding of CIs, yielding average gains of 1.7%, 2.0%, and 2.9% across eleven benchmarks, respectively. Satya Narayan Shukla, Mahmoud Azab, Aashu Singh, Qifan Wang 0001, Shengyun Peng, Hanchao Yu, Shen Yan 0007, Xuewen Zhang, Baosheng He |
ICCV | 2 |
| 2024 | The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language VariantsabstractLucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, Madian Khabsa. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal 0001, Abhinandan Krishnan, Luke Zettlemoyer, Madian Khabsa |
ACL (1) | 5 |
| 2024 | Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMsabstractIntegration of Large Language Models (LLMs) into visual domain tasks, resulting in visual-LLMs (V-LLMs), has enabled exceptional performance in vision-language tasks, particularly for visual question answering (VQA). However, existing V-LLMs (e.g. BLIP-2, LLaVA) demonstrate weak spatial reasoning and localization awareness. Despite generating highly descriptive and elaborate textual answers, these models fail at simple tasks like distinguishing a left vs right location. In this work, we explore how image-space coordinate based instruction fine-tuning objectives could inject spatial awareness into V-LLMs. We discover optimal coordinate representations, data-efficient instruction fine-tuning objectives, and pseudo-data generation strategies that lead to improved spatial awareness in V-LLMs. Additionally, our resulting model improves VQA across image and video domains, reduces undesired hallucination, and generates better contextual object descriptions. Experiments across 5 vision-language tasks involving 14 different datasets establish the clear performance improvements achieved by our proposed framework. Kanchana Ranasinghe, Satya Narayan Shukla, Omid Poursaeed, Michael S. Ryoo, Tsung-Yu Lin |
CVPR | 2 |
| 2024 | uCAP: An Unsupervised Prompting Method for Vision-Language Models
A. Tuan Nguyen, Kai Sheng Tai, Bor-Chun Chen, Satya Narayan Shukla, Hanchao Yu, Philip Torr 0001, Tai-Peng Tian, Ser-Nam Lim |
ECCV (74) | 4 |
| 2022 | Heteroscedastic Temporal Variational Autoencoder For Irregularly Sampled Time Series
Satya Narayan Shukla, Benjamin M. Marlin |
ICLR | 1 |
| 2021 | Multi-Time Attention Networks for Irregularly Sampled Time Series
Satya Narayan Shukla, Benjamin M. Marlin |
ICLR | 1 |
| 2021 | Simple and Efficient Hard Label Black-box Adversarial Attacks in Low Query Budget RegimesabstractWe focus on the problem of black-box adversarial attacks, where the aim is to generate adversarial examples for deep learning models solely based on information limited to output label (hard-label) to a queried data input. We propose a simple and efficient Bayesian Optimization (BO) based approach for developing black-box adversarial attacks. Issues with BO's performance in high dimensions are avoided by searching for adversarial examples in a structured low-dimensional subspace. We demonstrate the efficacy of our proposed attack method by evaluating both ℓ∞ and ℓ2 norm constrained untargeted and targeted hard label black-box attacks on three standard datasets - MNIST, CIFAR-10, and ImageNet. Our proposed approach consistently achieves 2x to 10x higher attack success rate while requiring 10x to 20x fewer queries compared to the current state-of-the-art black-box adversarial attacks. Satya Narayan Shukla, Anit Kumar Sahu, Devin Willmott, J. Zico Kolter |
KDD | 1 |
| 2019 | Interpolation-Prediction Networks for Irregularly Sampled Time Series
Satya Narayan Shukla, Benjamin M. Marlin |
ICLR (Poster) | 1 |