VLDB 2026 Research / reviewers in the wild / expert
Ishan Kavathekar
dblp:369/3652
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0005-8705-8305ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Trustworthy machine learning · 75% Multi-agent systems · 25% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › robustness
adversarial attack |
1.0 | 1 | 2026 | TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems · ACL (1) 2026 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
1.0 | 1 | 2026 | TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems · ACL (1) 2026 |
Knowledge, reasoning and agents › Multi-agent systems
LLM-based multi-agent systems |
1.0 | 1 | 2026 | TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems · ACL (1) 2026 |
Machine learning › Trustworthy machine learning
robustness |
1.0 | 1 | 2026 | TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems · ACL (1) 2026 |
Methods — techniques the papers use, named apart from their topics
benchmarking · 1.0adversarial attack · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM SystemsabstractLarge Language Models (LLMs) have demonstrated strong capabilities as autonomous agents through tool use, planning, and decisionmaking abilities, leading to their widespread adoption across diverse tasks.As task complexity grows, multi-agent LLM systems are increasingly used to solve problems collaboratively.However, safety and security of these systems remains largely under-explored.Existing benchmarks and datasets predominantly focus on single-agent settings, failing to capture the unique vulnerabilities of multi-agent dynamics and co-ordination.To address this gap, we introduce Threats and Attacks in Multi-Agent Systems (TAMAS), a benchmark designed to evaluate the robustness and safety of multi-agent LLM systems.TAMAS includes five distinct scenarios comprising 300 adversarial instances across six attack types and 211 tools, along with 100 harmless tasks.We assess system performance across ten backbone LLMs and three agent interaction configurations from Autogen and CrewAI frameworks, highlighting critical challenges and failure modes in current multi-agent deployments.Furthermore, we introduce Effective Robustness Score (ERS) to assess the tradeoff between safety and task effectiveness of these frameworks.Our findings show that multi-agent systems are highly vulnerable to adversarial attacks, underscoring the urgent need for stronger defenses.TAMAS provides a foundation for systematically studying and improving the safety of multi-agent LLM systems.Code and dataset is available at https://github.com/microsoft/TAMAS. Ishan Kavathekar, Hemang Jain, Ameya Rathod, Ponnurangam Kumaraguru, Tanuja Ganu |
ACL (1) | 1 |
| 2025 | Small Models, Big Tasks: An Exploratory Empirical Study on Small Language Models for Function CallingabstractFunction calling is a complex task with widespread applications in domains such as information retrieval, software engineering and automation. For example, a query to book the shortest flight from New York to London on January 15 requires identifying the correct parameters to generate accurate function calls. Large Language Models (LLMs) can automate this process but are computationally expensive and impractical in resource-constrained settings. In contrast, Small Language Models (SLMs) can operate efficiently, offering faster response times, and lower computational demands, making them potential candidates for function calling on edge devices. In this exploratory empirical study, we evaluate the efficacy of SLMs in generating function calls across diverse domains using zero-shot, few-shot, and fine-tuning approaches, both with and without prompt injection, while also providing the finetuned models to facilitate future applications. Furthermore, we analyze the model responses across a range of metrics, capturing various aspects of function call generation. Additionally, we perform experiments on an edge device to evaluate their performance in terms of latency and memory usage, providing useful insights into their practical applicability. Our findings show that while SLMs improve from zero-shot to few-shot and perform best with fine-tuning, they struggle significantly with adhering to the given output format. Prompt injection experiments further indicate that the models are generally robust and exhibit only a slight decline in performance. While SLMs demonstrate potential for the function call generation task, our results also highlight areas that need further refinement for real-time functioning. Ishan Kavathekar, Raghav Donakanti, Ponnurangam Kumaraguru, Karthik Vaidhyanathan |
EASE | 1 |
| 2024 | InSaAF: Incorporating Safety Through Accuracy and Fairness - Are LLMs Ready for the Indian Legal Domain?abstractLarge Language Models (LLMs) have emerged as powerful tools to perform various tasks in the legal domain, ranging from generating summaries to predicting judgments. Despite their immense potential, these models have been proven to learn and exhibit societal biases and make unfair predictions. Hence, it is essential to evaluate these models prior to deployment. In this study, we explore the ability of LLMs to perform Binary Statutory Reasoning in the Indian legal landscape across various societal disparities. We present a novel metric, β-weighted Legal Safety Score (LSSβ), to evaluate the legal usability of the LLMs. Additionally, we propose a finetuning pipeline, utilising specialised legal datasets, as a potential method to reduce bias. Our proposed pipeline effectively reduces bias in the model, as indicated by improved LSSβ. This highlights the potential of our approach to enhance fairness in LLMs, making them more reliable for legal tasks in socially diverse contexts. Yogesh Tripathi, Raghav Donakanti, Sahil Girhepuje, Ishan Kavathekar, Bhaskara Hanuma Vedula, Gokul S. Krishnan, Anmol Goel, Shreya Goyal, Balaraman Ravindran, Ponnurangam Kumaraguru |
JURIX | 4 |