Robert Osazuwa Ness

dblp:198/4937 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Trustworthy machine learning · 53% Probabilistic and Bayesian machine learning · 27% Language models and text generation · 12%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 86% Computational science and engineering · 14%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
0.912025
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations · ICLR 2025
Machine learning › Trustworthy machine learning › language model interpretability
large language model explanation
0.912025
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations · ICLR 2025
Machine learning › Trustworthy machine learning
cognitive evaluation
0.712023
Evaluating Cognitive Maps and Planning in Large Language Models with CogEval · NeurIPS 2023
Natural language and speech › Language models and text generation
large language model evaluation
0.712023
Evaluating Cognitive Maps and Planning in Large Language Models with CogEval · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.412019
Integrating Markov processes with structural causal modeling enables counterfactual inference in complex systems · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › causal inference
counterfactual prediction
0.412019
Integrating Markov processes with structural causal modeling enables counterfactual inference in complex systems · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
markov processes
0.412019
Integrating Markov processes with structural causal modeling enables counterfactual inference in complex systems · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal model
structural causal model
0.412019
Integrating Markov processes with structural causal modeling enables counterfactual inference in complex systems · NeurIPS 2019
Bioinformatics and computational biology › biological network › network biology
signaling network inference
0.312017
A Bayesian Active Learning Experimental Design for Inferring Signaling Networks · RECOMB 2017
Machine learning › Trustworthy machine learning
fairness
0.312025
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations · ICLR 2025
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
medical question answering
0.312025
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations · ICLR 2025
Machine learning › Trustworthy machine learning › fairness
social bias
0.312025
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations · ICLR 2025
Robotics › Robot navigation and mapping › robot mapping
cognitive map
0.212023
Evaluating Cognitive Maps and Planning in Large Language Models with CogEval · NeurIPS 2023
Bioinformatics and computational biology › systems biology › parameter estimation
identifiability analysis
0.212022
Do-calculus enables estimation of causal effects in partially observed biomolecular pathways · Bioinform. 2022
Computational science and engineering
latent variable model
0.212022
Do-calculus enables estimation of causal effects in partially observed biomolecular pathways · Bioinform. 2022

Methods — techniques the papers use, named apart from their topics

hierarchical bayesian modeling · 0.9counterfactual generation · 0.9causal effect estimation · 0.9statistical robustness tests · 0.7cognitive science-inspired protocol · 0.7latent variable model · 0.6do-calculus · 0.6structural causal modeling · 0.4markov process modeling · 0.4bayesian experimental design · 0.3active learning · 0.3
YearPublicationVenuePosition
2025 Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
abstract
Large language models (LLMs) are capable of generating *plausible* explanations of how they arrived at an answer to a question. However, these explanations can misrepresent the model's "reasoning" process, i.e., they can be *unfaithful*. This, in turn, can lead to over-trust and misuse. We introduce a new approach for measuring the faithfulness of LLM explanations. First, we provide a rigorous definition of faithfulness. Since LLM explanations mimic human explanations, they often reference high-level *concepts* in the input question that purportedly influenced the model. We define faithfulness in terms of the difference between the set of concepts that the LLM's *explanations imply* are influential and the set that *truly* are. Second, we present a novel method for estimating faithfulness that is based on: (1) using an auxiliary LLM to modify the values of concepts within model inputs to create realistic counterfactuals, and (2) using a hierarchical Bayesian model to quantify the causal effects of concepts at both the example- and dataset-level. Our experiments show that our method can be used to quantify and discover interpretable patterns of unfaithfulness. On a social bias task, we uncover cases where LLM explanations hide the influence of social bias. On a medical question answering task, we uncover cases where LLM explanations provide misleading claims about which pieces of evidence influenced the model's decisions.
Katie Matton, Robert Osazuwa Ness, John V. Guttag, Emre Kiciman
ICLR2
2023 Evaluating Cognitive Maps and Planning in Large Language Models with CogEval
abstract
Recently an influx of studies claims emergent cognitive abilities in large language models (LLMs). Yet, most rely on anecdotes, overlook contamination of training sets, or lack systematic Evaluation involving multiple tasks, control conditions, multiple iterations, and statistical robustness tests. Here we make two major contributions. First, we propose CogEval, a cognitive science-inspired protocol for the systematic evaluation of cognitive capacities in LLMs. The CogEval protocol can be followed for the evaluation of various abilities. Second, here we follow CogEval to systematically evaluate cognitive maps and planning ability across eight LLMs (OpenAI GPT-4, GPT-3.5-turbo-175B, davinci-003-175B, Google Bard, Cohere-xlarge-52.4B, Anthropic Claude-1-52B, LLaMA-13B, and Alpaca-7B). We base our task prompts on human experiments, which offer both established construct validity for evaluating planning, and are absent from LLM training sets. We find that, while LLMs show apparent competence in a few planning tasks with simpler structures, systematic evaluation reveals striking failure modes in planning tasks, including hallucinations of invalid trajectories and falling in loops. These findings do not support the idea of emergent out-of-the-box planning ability in LLMs. This could be because LLMs do not understand the latent relational structures underlying planning problems, known as cognitive maps, and fail at unrolling goal-directed trajectories based on the underlying structure. Implications for application and future directions are discussed.
Ida Momennejad, Hosein Hasanbeig, Felipe Vieira Frujeri, Hiteshi Sharma, Nebojsa Jojic, Hamid Palangi, Robert Osazuwa Ness, Jonathan Larson
NeurIPS7
2022 Do-calculus enables estimation of causal effects in partially observed biomolecular pathways
abstract
MOTIVATION: Estimating causal queries, such as changes in protein abundance in response to a perturbation, is a fundamental task in the analysis of biomolecular pathways. The estimation requires experimental measurements on the pathway components. However, in practice many pathway components are left unobserved (latent) because they are either unknown, or difficult to measure. Latent variable models (LVMs) are well-suited for such estimation. Unfortunately, LVM-based estimation of causal queries can be inaccurate when parameters of the latent variables are not uniquely identified, or when the number of latent variables is misspecified. This has limited the use of LVMs for causal inference in biomolecular pathways. RESULTS: In this article, we propose a general and practical approach for LVM-based estimation of causal queries. We prove that, despite the challenges above, LVM-based estimators of causal queries are accurate if the queries are identifiable according to Pearl's do-calculus and describe an algorithm for its estimation. We illustrate the breadth and the practical utility of this approach for estimating causal queries in four synthetic and two experimental case studies, where structures of biomolecular pathways challenge the existing methods for causal query estimation. AVAILABILITY AND IMPLEMENTATION: The code and the data documenting all the case studies are available at https://github.com/srtaheri/LVMwithDoCalculus. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sara Mohammad Taheri, Jeremy Zucker, Charles Tapley Hoyt, Karen Sachs, Vartika Tewari, Robert Osazuwa Ness, Olga Vitek
Bioinform.6
2021 Leveraging Structured Biological Knowledge for Counterfactual Inference: A Case Study of Viral Pathogenesis
abstract
Counterfactual inference is a useful tool for comparing outcomes of interventions on complex systems. It requires us to represent the system in form of a structural causal model, complete with a causal diagram, probabilistic assumptions on exogenous variables, and functional assignments. Specifying such models can be extremely difficult in practice. The process requires substantial domain expertise, and does not scale easily to large systems, multiple systems, or novel system modifications. At the same time, many application domains, such as molecular biology, are rich in structured causal knowledge that is qualitative in nature. This article proposes a general approach for querying a causal biological knowledge graph, and converting the qualitative result into a quantitative structural causal model that can learn from data to answer the question. We demonstrate the feasibility, accuracy and versatility of this approach using two case studies in systems biology. The first demonstrates the appropriateness of the underlying assumptions and the accuracy of the results. The second demonstrates the versatility of the approach by querying a knowledge base for the molecular determinants of a severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2)-induced cytokine storm, and performing counterfactual inference to estimate the causal effect of medical countermeasures for severely ill patients.
Jeremy Zucker, Kaushal Paneri, Sara Mohammad Taheri, Somya Bhargava, Pallavi Kolambkar, Craig Bakker, Jeremy Teuton, Charles Tapley Hoyt, Kristie L. Oxford, Robert Osazuwa Ness, Olga Vitek
IEEE Trans. Big Data10
2019 Integrating Markov processes with structural causal modeling enables counterfactual inference in complex systems
abstract
This manuscript contributes a general and practical framework for casting a Markov process model of a system at equilibrium as a structural causal model, and carrying out counterfactual inference. Markov processes mathematically describe the mechanisms in the system, and predict the system’s equilibrium behavior upon intervention, but do not support counterfactual inference. In contrast, structural causal models support counterfactual inference, but do not identify the mechanisms. This manuscript leverages the benefits of both approaches. We define the structural causal models in terms of the parameters and the equilibrium dynamics of the Markov process models, and counterfactual inference flows from these settings. The proposed approach alleviates the identifiability drawback of the structural causal models, in that the counterfactual inference is consistent with the counterfactual trajectories simulated from the Markov process model. We showcase the benefits of this framework in case studies of complex biomolecular systems with nonlinear dynamics. We illustrate that, in presence of Markov process model misspecification, counterfactual inference leverages prior data, and therefore estimates the outcome of an intervention more accurately than a direct simulation.
Robert Osazuwa Ness, Kaushal Paneri, Olga Vitek
NeurIPS1
2017 A Bayesian Active Learning Experimental Design for Inferring Signaling Networks
Robert Osazuwa Ness, Karen Sachs, Parag Mallick, Olga Vitek
RECOMB1