Ankur Taly

dblp:60/3530 · DBLP profile ↗
← Back
24ranked-venue papers
5as first author
5since 2021 · last 2025
0009-0004-3459-0288ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 since 2021Security and privacy · 6 · 1 first-authorSoftware engineering, systems software and programming languages · 6 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Trustworthy machine learning · 50% Language models and text generation · 31% Efficient and distributed learning · 8%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%
Network and information security
3 papers
Authentication and access control · 47% Web and mobile security · 39% Systems and software security · 14%
Software engineering, system software, and programming languages
4 papers
Program analysis · 52% Programming languages and type systems · 27% Program verification · 21%

Topics — the 30 heaviest of 35, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
1.952022
First is Better Than Last for Language Data Influence · NeurIPS 2022
Explainable AI in Industry · KDD 2019
Property Inference for Deep Neural Networks · ASE 2019
Natural language and speech › Language models and text generation
retrieval-augmented generation
1.722025
Sufficient Context: A New Lens on Retrieval Augmented Generation Systems · ICLR 2025
Speculative RAG: Enhancing Retrieval Augmented Generation through Drafting · ICLR 2025
Natural language and speech › Language models and text generation
hallucination mitigation
0.912025
Sufficient Context: A New Lens on Retrieval Augmented Generation Systems · ICLR 2025
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding
0.912025
Speculative RAG: Enhancing Retrieval Augmented Generation through Drafting · ICLR 2025
Information retrieval › evaluation › text generation evaluation
retrieval-augmented generation evaluation
0.912025
Sufficient Context: A New Lens on Retrieval Augmented Generation Systems · ICLR 2025
Machine learning › Trustworthy machine learning
hallucination
0.812024
Grounding and Evaluation for Large Language Models: Practical Challenges and Lessons Learned (Survey) · KDD 2024
Natural language and speech › Language models and text generation
large language model evaluation
0.812024
Grounding and Evaluation for Large Language Models: Practical Challenges and Lessons Learned (Survey) · KDD 2024
Machine learning › Trustworthy machine learning › interpretability
training data attribution
0.612022
First is Better Than Last for Language Data Influence · NeurIPS 2022
Machine learning › Trustworthy machine learning › interpretability
explainable AI
0.412019
Explainable AI in Industry · KDD 2019
Machine learning › Trustworthy machine learning
robustness
0.412019
Property Inference for Deep Neural Networks · ASE 2019
Machine learning › Trustworthy machine learning › robustness
robustness guarantees
0.412019
Property Inference for Deep Neural Networks · ASE 2019
Machine learning › Trustworthy machine learning › interpretability
attribution methods
0.312018
Did the Model Understand the Question? · ACL (1) 2018
Natural language and speech › Question answering and dialogue systems
machine reading comprehension
0.312018
Did the Model Understand the Question? · ACL (1) 2018
Natural language and speech › Question answering and dialogue systems
table question answering
0.312018
Did the Model Understand the Question? · ACL (1) 2018
Computer vision › Vision and language
visual question answering
0.312018
Did the Model Understand the Question? · ACL (1) 2018
Machine learning › Trustworthy machine learning › interpretability › attribution methods
axiomatic attribution
0.312017
Axiomatic Attribution for Deep Networks · ICML 2017
Machine learning › Trustworthy machine learning › interpretability › attribution methods
feature attribution
0.312017
Axiomatic Attribution for Deep Networks · ICML 2017
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation
grounding
0.212024
Grounding and Evaluation for Large Language Models: Practical Challenges and Lessons Learned (Survey) · KDD 2024
Authentication and access control
authorization
0.212014
Macaroons: Cookies with Contextual Caveats for Decentralized Authorization in the Cloud · NDSS 2014
Authentication and access control › authorization
decentralized authorization
0.212014
Macaroons: Cookies with Contextual Caveats for Decentralized Authorization in the Cloud · NDSS 2014
Program analysis › binary analysis
static binary analysis
0.112012
Automated synthesis of symbolic instruction encodings from I/O samples · PLDI 2012
Program analysis
symbolic execution
0.112012
Automated synthesis of symbolic instruction encodings from I/O samples · PLDI 2012
Web and mobile security › web security
API security
0.112011
Automated Analysis of Security-Critical JavaScript APIs · IEEE Symposium on Security and Privacy 2011
Web and mobile security
javascript security
0.112011
Automated Analysis of Security-Critical JavaScript APIs · IEEE Symposium on Security and Privacy 2011
Machine learning › Trustworthy machine learning
fairness
0.112019
Explainable AI in Industry · KDD 2019
Program verification
neural network verification
0.112019
Property Inference for Deep Neural Networks · ASE 2019
Systems and software security › isolation
isolation of untrusted code
0.112010
Object Capabilities and Isolation of Untrusted Web Applications · IEEE Symposium on Security and Privacy 2010
Programming languages and type systems
language-based security
0.112010
Object Capabilities and Isolation of Untrusted Web Applications · IEEE Symposium on Security and Privacy 2010
Cloud and datacenter computing › cloud security
cloud access control
0.112014
Macaroons: Cookies with Contextual Caveats for Decentralized Authorization in the Cloud · NDSS 2014
Programming languages and type systems
language semantics
0.012011
Automated Analysis of Security-Critical JavaScript APIs · IEEE Symposium on Security and Privacy 2011

Methods — techniques the papers use, named apart from their topics

verification · 1.7selective generation · 1.7drafting · 1.7distillation · 1.7context sufficiency classification · 1.7monitoring · 0.8evaluation · 0.8word embedding layer · 0.6gradient-based influence · 0.6adversarial perturbation · 0.3formal semantics · 0.2automated analysis · 0.2language-based isolation proofs · 0.2authority safety · 0.2smart sampling · 0.1bit-vector constraint synthesis · 0.1SMT solving · 0.1
YearPublicationVenuePosition
2025 Speculative RAG: Enhancing Retrieval Augmented Generation through Drafting
abstract
Retrieval augmented generation (RAG) combines the generative abilities of large language models (LLMs) with external knowledge sources to provide more accurate and up-to-date responses. Recent RAG advancements focus on improving retrieval outcomes through iterative LLM refinement or self-critique capabilities acquired through additional instruction tuning of LLMs. In this work, we introduce Speculative RAG - a framework that leverages a larger generalist LM to efficiently verify multiple RAG drafts produced in parallel by a smaller, distilled specialist LM. Each draft is generated from a distinct subset of retrieved documents, offering diverse perspectives on the evidence while reducing input token counts per draft. This approach enhances comprehension of each subset and mitigates potential position bias over long context. Our method accelerates RAG by delegating drafting to the smaller specialist LM, with the larger generalist LM performing a single verification pass over the drafts. Extensive experiments demonstrate that Speculative RAG achieves state-of-the-art performance with reduced latency on TriviaQA, MuSiQue, PopQA, PubHealth, and ARC-Challenge benchmarks. It notably enhances accuracy by up to 12.97% while reducing latency by 50.83% compared to conventional RAG systems on PubHealth.
Zilong Wang 0002, Zifeng Wang 0002, Long T. Le, Huaixiu Steven Zheng, Swaroop Mishra, Vincent Perot, Yuwei Zhang 0001, Anush Mattapalli, Ankur Taly, Jingbo Shang, Chen-Yu Lee, Tomas Pfister
ICLR9
2025 Sufficient Context: A New Lens on Retrieval Augmented Generation Systems
abstract
Augmenting LLMs with context leads to improved performance across many applications. Despite much research on Retrieval Augmented Generation (RAG) systems, an open question is whether errors arise because LLMs fail to utilize the context from retrieval or the context itself is insufficient to answer the query. To shed light on this, we develop a new notion of sufficient context, along with a method to classify instances that have enough information to answer the query. We then use sufficient context to analyze several models and datasets. By stratifying errors based on context sufficiency, we find that larger models with higher baseline performance (Gemini 1.5 Pro, GPT 4o, Claude 3.5) excel at answering queries when the context is sufficient, but often output incorrect answers instead of abstaining when the context is not. On the other hand, smaller models with lower baseline performance (Llama 3.1, Mistral 3, Gemma 2) hallucinate or abstain often, even with sufficient context. We further categorize cases when the context is useful, and improves accuracy, even though it does not fully answer the query and the model errs without the context. Building on our findings, we explore ways to reduce hallucinations in RAG systems, including a new selective generation method that leverages sufficient context information for guided abstention. Our method improves the fraction of correct answers among times where the model responds by 2--10% for Gemini, GPT, and Gemma. Code for our selective generation method and the prompts used in our autorater analysis are available on our [github](https://github.com/hljoren/sufficientcontext).
Hailey Joren, Chun-Sung Ferng, Da-Cheng Juan, Ankur Taly, Cyrus Rashtchian
ICLR5
2024 Grounding and Evaluation for Large Language Models: Practical Challenges and Lessons Learned (Survey)
abstract
With the ongoing rapid adoption of Artificial Intelligence (AI)-based systems in high-stakes domains, ensuring the trustworthiness, safety, and observability of these systems has become crucial. It is essential to evaluate and monitor AI systems not only for accuracy and quality-related metrics but also for robustness, bias, security, interpretability, and other responsible AI dimensions. We focus on large language models (LLMs) and other generative AI models, which present additional challenges such as hallucinations, harmful and manipulative content, and copyright infringement. In this survey article accompanying our tutorial, we highlight a wide range of harms associated with generative AI systems, and survey state of the art approaches (along with open challenges) to address these harms.
Krishnaram Kenthapadi, Mehrnoosh Sameki, Ankur Taly
KDD3
2022 First is Better Than Last for Language Data Influence
abstract
The ability to identify influential training examples enables us to debug training data and explain model behavior. Existing techniques to do so are based on the flow of training data influence through the model parameters. For large models in NLP applications, it is often computationally infeasible to study this flow through all model parameters, therefore techniques usually pick the last layer of weights. However, we observe that since the activation connected to the last layer of weights contains "shared logic", the data influenced calculated via the last layer weights prone to a "cancellation effect", where the data influence of different examples have large magnitude that contradicts each other. The cancellation effect lowers the discriminative power of the influence score, and deleting influential examples according to this measure often does not change the model's behavior by much. To mitigate this, we propose a technique called TracIn-WE that modifies a method called TracIn to operate on the word embedding layer instead of the last layer, where the cancellation effect is less severe. One potential concern is that influence based on the word embedding layer may not encode sufficient high level information. However, we find that gradients (unlike embeddings) do not suffer from this, possibly because they chain through higher layers. We show that TracIn-WE significantly outperforms other data influence methods applied on the last layer significantly on the case deletion evaluation on three language classification tasks for different models. In addition, TracIn-WE can produce scores not just at the level of the overall training input, but also at the level of words within the training input, a further aid in debugging.
Chih-Kuan Yeh, Ankur Taly, Mukund Sundararajan, Frederick Liu, Pradeep Ravikumar
NeurIPS2
2021 Local explanations via necessity and sufficiency: unifying theory and practice
abstract
Necessity and sufficiency are the building blocks of all successful explanations. Yet despite their importance, these notions have been conceptually underdeveloped and inconsistently applied in explainable artificial intelligence (XAI), a fast-growing research area that is so far lacking in firm theoretical foundations. Building on work in logic, probability, and causality, we establish the central role of necessity and sufficiency in XAI, unifying seemingly disparate methods in a single formal framework. We provide a sound and complete algorithm for computing explanatory factors with respect to a given context, and demonstrate its flexibility and competitive performance against state of the art alternatives on various tasks.
David S. Watson, Limor Gultchin, Ankur Taly, Luciano Floridi
UAI3
2020 The Explanation Game: Explaining Machine Learning Models Using Shapley Values
Luke Merrick, Ankur Taly
CD-MAKE2
2019 Counterfactual Fairness in Text Classification through Robustness
abstract
In this paper, we study counterfactual fairness in text classification, which asks the question: How would the prediction change if the sensitive attribute referenced in the example were different? Toxicity classifiers demonstrate a counterfactual fairness issue by predicting that "Some people are gay" is toxic while "Some people are straight" is nontoxic. We offer a metric, counterfactual token fairness (CTF), for measuring this particular form of fairness in text classifiers, and describe its relationship with group fairness. Further, we offer three approaches, blindness, counterfactual augmentation, and counterfactual logit pairing (CLP), for optimizing counterfactual token fairness during training, bridging the robustness and fairness literature. Empirically, we find that blindness and CLP address counterfactual token fairness. The methods do not harm classifier performance, and have varying tradeoffs with group fairness. These approaches, both for measurement and optimization, provide a new path forward for addressing fairness concerns in text classification.
Sahaj Garg, Vincent Perot, Nicole Limtiaco, Ankur Taly, Ed H. Chi, Alex Beutel
AIES4
2019 Property Inference for Deep Neural Networks
abstract
We present techniques for automatically inferring formal properties of feed-forward neural networks. We observe that a significant part (if not all) of the logic of feed forward networks is captured in the activation status (on or off) of its neurons. We propose to extract patterns based on neuron decisions as preconditions that imply certain desirable output property e.g., the prediction being a certain class. We present techniques to extract input properties, encoding convex predicates on the input space that imply given output properties and layer properties, representing network properties captured in the hidden layers that imply the desired output behavior. We apply our techniques on networks for the MNIST and ACASXU applications. Our experiments highlight the use of the inferred properties in a variety of tasks, such as explaining predictions, providing robustness guarantees, simplifying proofs, and network distillation.
Divya Gopinath, Hayes Converse, Corina Pasareanu, Ankur Taly
ASE4
2019 Explainable AI in Industry
abstract
Artificial Intelligence is increasingly playing an integral role in determining our day-to-day experiences. Moreover, with proliferation of AI based solutions in areas such as hiring, lending, criminal justice, healthcare, and education, the resulting personal and professional implications of AI are far-reaching. The dominant role played by AI models in these domains has led to a growing concern regarding potential bias in these models, and a demand for model transparency and interpretability. In addition, model explainability is a prerequisite for building trust and adoption of AI systems in high stakes domains requiring reliability and safety such as healthcare and automated transportation, and critical industrial applications with significant economic implications such as predictive maintenance, exploration of natural resources, and climate change modeling.
Krishna Gade, Sahin Cem Geyik, Krishnaram Kenthapadi, Varun Mithal, Ankur Taly
KDD5
2018 Did the Model Understand the Question?
abstract
We analyze state-of-the-art deep learning models for three tasks: question answering on (1) images, (2) tables, and (3) passages of text.Using the notion of attribution (word importance), we find that these deep networks often ignore important question terms.Leveraging such behavior, we perturb questions to craft a variety of adversarial examples.Our strongest attacks drop the accuracy of a visual question answering model from 61.1% to 19%, and that of a tabular question answering model from 33.5% to 3.3%.Additionally, we show how attributions can strengthen attacks proposed by Jia and Liang (2017) on paragraph comprehension models.Our results demonstrate that attributions can augment standard measures of accuracy and empower investigation of model performance.When a model is accurate but for the wrong reasons, attributions can surface erroneous logic in the model that indicates inadequacies in the test data.
Pramod Kaushik Mudrakarta, Ankur Taly, Mukund Sundararajan, Kedar Dhamdhere
ACL (1)2
2017 Axiomatic Attribution for Deep Networks
abstract
We study the problem of attributing the prediction of a deep network to its input features, a problem previously studied by several other works. We identify two fundamental axioms—Sensitivity and Implementation Invariance that attribution methods ought to satisfy. We show that they are not satisfied by most known attribution methods, which we consider to be a fundamental weakness of those methods. We use the axioms to guide the design of a new attribution method called Integrated Gradients. Our method requires no modification to the original network and is extremely simple to implement; it just needs a few calls to the standard gradient operator. We apply this method to a couple of image models, a couple of text models and a chemistry model, demonstrating its ability to debug networks, to extract rules from a network, and to enable users to engage with models better.
Mukund Sundararajan, Ankur Taly, Qiqi Yan
ICML2
2016 Privacy, Discovery, and Authentication for the Internet of Things
David J. Wu 0001, Ankur Taly, Asim Shankar, Dan Boneh
ESORICS (2)2
2014 Macaroons: Cookies with Contextual Caveats for Decentralized Authorization in the Cloud
Arnar Birgisson, Joe Gibbs Politz, Úlfar Erlingsson, Ankur Taly, Michael Vrable, Mark Lentczner
NDSS4
2012 Automated synthesis of symbolic instruction encodings from I/O samples
abstract
Symbolic execution is a key component of precise binary program analysis tools. We discuss how to automatically boot-strap the construction of a symbolic execution engine for a processor instruction set such as x86, x64 or ARM. We show how to automatically synthesize symbolic representations of individual processor instructions from input/output examples and express them as bit-vector constraints. We present and compare various synthesis algorithms and instruction sampling strategies. We introduce a new synthesis algorithm based on smart sampling which we show is one to two orders of magnitude faster than previous synthesis algorithms in our context. With this new algorithm, we can automatically synthesize bit-vector circuits for over 500 x86 instructions (8/16/32-bits, outputs, EFLAGS) using only 6 synthesis templates and in less than two hours using the Z3 SMT solver on a regular machine. During this work, we also discovered several inconsistencies across x86 processors, errors in the x86 Intel spec, and several bugs in previous manually-written x86 instruction handlers.
Patrice Godefroid, Ankur Taly
PLDI2
2011 Automated Analysis of Security-Critical JavaScript APIs
abstract
JavaScript is widely used to provide client-side functionality in Web applications. To provide services ranging from maps to advertisements, Web applications may incorporate untrusted JavaScript code from third parties. The trusted portion of each application may then expose an API to untrusted code, interposing a reference monitor that mediates access to security-critical resources. However, a JavaScript reference monitor can only be effective if it cannot be circumvented through programming tricks or programming language idiosyncrasies. In order to verify complete mediation of critical resources for applications of interest, we define the semantics of a restricted version of JavaScript devised by the ECMA Standards committee for isolation purposes, and develop and test an automated tool that can soundly establish that a given API cannot be circumvented or subverted. Our tool reveals a previously-undiscovered vulnerability in the widely-examined Yahoo! AD Safe filter and verifies confinement of the repaired filter and other examples from the Object-Capability literature.
Ankur Taly, Úlfar Erlingsson, John C. Mitchell, Mark S. Miller, Jasvir Nagra
IEEE Symposium on Security and Privacy1
2011 Synthesizing switching logic using constraint solving
Ankur Taly, Sumit Gulwani, Ashish Tiwari 0001
Int. J. Softw. Tools Technol. Transf.1
2010 Switching logic synthesis for reachability
abstract
We consider the problem of driving a system from some initial configuration to a desired configuration while avoiding some unsafe configurations. The system to be controlled is a dynamical system that can operate in different modes. The goal is to synthesize the logic for switching between the modes so that the desired reachability property holds.
Ankur Taly, Ashish Tiwari 0001
EMSOFT1
2010 Object Capabilities and Isolation of Untrusted Web Applications
abstract
A growing number of current web sites combine active content (applications) from untrusted sources, as in so-called mashups. The object-capability model provides an appealing approach for isolating untrusted content: if separate applications are provided disjoint capabilities, a sound object capability framework should prevent untrusted applications from interfering with each other, without preventing interaction with the user or the hosting page. In developing language-based foundations for isolation proofs based on object-capability concepts, we identify a more general notion of authority safety that also implies resource isolation. After proving that capability safety implies authority safety, we show the applicability of our framework for a specific class of mashups. In addition to proving that a JavaScript subset based on Google Caja is capability safe, we prove that a more expressive subset of JavaScript is authority safe, even though it is not based on the object-capability model.
Sergio Maffeis, John C. Mitchell, Ankur Taly
IEEE Symposium on Security and Privacy3
2009 Language-Based Isolation of Untrusted JavaScript
abstract
Web sites that incorporate untrusted content may use browser- or language-based methods to keep such content from maliciously altering pages, stealing sensitive information, or causing other harm. We study language-based methods for filtering and rewriting JavaScript code, using Yahoo! ADSafe and Facebook FBJS as motivating examples. We explain the core problems by describing previously unknown vulnerabilities and subtleties, and develop a foundation for improved solutions based on an operational semantics of the full ECMA-262 language. We also discuss how to apply our analysis to address the JavaScript isolation problems we discovered.
Sergio Maffeis, Ankur Taly
CSF2
2009 Isolating JavaScript with Filters, Rewriting, and Wrappers
Sergio Maffeis, John C. Mitchell, Ankur Taly
ESORICS3
2009 Deductive Verification of Continuous Dynamical Systems
abstract
We define the notion of inductive invariants for continuous dynamical systems and use it to present inference rules for safety verification of polynomial continuous dynamical systems. We present two different sound and complete inference rules, but neither of these rules can be effectively applied. We then present several simpler and practical inference rules that are sound and relatively complete for different classes of inductive invariants. The simpler inference rules can be effectively checked when all involved sets are semi-algebraic.
Ankur Taly, Ashish Tiwari 0001
FSTTCS1
2009 Synthesizing Switching Logic Using Constraint Solving
Ankur Taly, Sumit Gulwani, Ashish Tiwari 0001
VMCAI1
2008 An Operational Semantics for JavaScript
Sergio Maffeis, John C. Mitchell, Ankur Taly
APLAS3
2007 Static Analysis by Policy Iteration on Relational Domains
Stéphane Gaubert, Eric Goubault, Ankur Taly, Sarah Zennou
ESOP3