Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Christopher Mohri

dblp:230/3446 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Trustworthy machine learning · 36% Learning theory · 28% Language models and text generation · 21%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › uncertainty estimation
conformal prediction
1.622025
Online Conformal Prediction via Online Optimization · ICML 2025
Language Models with Conformal Factuality Guarantees · ICML 2024
Machine learning › Trustworthy machine learning
uncertainty estimation
1.622025
Online Conformal Prediction via Online Optimization · ICML 2025
Language Models with Conformal Factuality Guarantees · ICML 2024
Machine learning › Learning theory › excess risk bounds
h-consistency bounds
1.422024
Cardinality-Aware Set Prediction and Top-$k$ Classification · NeurIPS 2024
Two-Stage Learning to Defer with Multiple Experts · NeurIPS 2023
Natural language and speech › Question answering and dialogue systems
answer aggregation
1.012026
Algorithmic Thinking Theory · COLT 2026
Natural language and speech › Language models and text generation
large language model reasoning
1.012026
Algorithmic Thinking Theory · COLT 2026
Machine learning › Learning theory › loss function
surrogate loss
1.022024
Cardinality-Aware Set Prediction and Top-$k$ Classification · NeurIPS 2024
Learning to Reject with a Fixed Predictor: Application to Decontextualization · ICLR 2024
Mathematical optimization
online optimization
0.912025
Online Conformal Prediction via Online Optimization · ICML 2025
Natural language and speech › Language models and text generation › natural language understanding
decontextualization
0.812024
Learning to Reject with a Fixed Predictor: Application to Decontextualization · ICLR 2024
Machine learning › Trustworthy machine learning › uncertainty estimation
selective classification
0.812024
Learning to Reject with a Fixed Predictor: Application to Decontextualization · ICLR 2024
Machine learning › Learning theory › statistical estimation › consistency analysis
consistency guarantees
0.712023
Two-Stage Learning to Defer with Multiple Experts · NeurIPS 2023
Knowledge, reasoning and agents › Multi-agent systems › human-agent interaction › human-AI decision making
learning to defer
0.712023
Two-Stage Learning to Defer with Multiple Experts · NeurIPS 2023
Machine learning › Trustworthy machine learning › uncertainty estimation › conformal prediction
coverage guarantee
0.312025
Online Conformal Prediction via Online Optimization · ICML 2025
Natural language and speech › Question answering and dialogue systems › knowledge-intensive question answering
closed-book question answering
0.212024
Language Models with Conformal Factuality Guarantees · ICML 2024
Machine learning › Learning theory › loss function › surrogate loss › consistency of surrogate losses
h-consistency
0.212024
Learning to Reject with a Fixed Predictor: Application to Decontextualization · ICLR 2024

Methods — techniques the papers use, named apart from their topics

quantile regression · 1.7online conformal prediction · 1.7probabilistic oracle model · 1.0surrogate loss minimization · 0.8cost-sensitive loss · 0.8conformal prediction · 0.8comp-sum loss · 0.8cross-entropy training · 0.7bayes-consistency analysis · 0.7
YearPublicationVenuePosition
2026 Algorithmic Thinking Theory
abstract
Large language models (LLMs) have proven to be highly effective for solving complex reasoning tasks. Surprisingly, their capabilities can often be improved by iterating on previously generated solutions. In this context, a reasoning plan for generating and combining a set of solutions can be thought of as an algorithm for reasoning using a probabilistic oracle. We introduce a theoretical framework for analyzing such reasoning algorithms. This framework formalizes the principles underlying popular techniques for iterative improvement and answer aggregation, providing a foundation for designing a new generation of more powerful reasoning methods. Unlike approaches for understanding models that rely on architectural specifics, our model is grounded in experimental evidence. As a result, it offers a general perspective that may extend to a wide range of current and future reasoning oracles.
Mohammad Hossein Bateni 0001, Vincent Cohen-Addad, Yuzhou Gu, Silvio Lattanzi, Simon Meierhans, Christopher Mohri
COLT6
2025 Online Conformal Prediction via Online Optimization
abstract
We introduce a family of algorithms for online conformal prediction with coverage guarantees for both adversarial and stochastic data. In the adversarial setting, we establish the standard guarantee: over time, a pre-specified target fraction of confidence sets cover the ground truth. For stochastic data, we provide a guarantee at every time instead of just on average over time: the probability that a confidence set covers the ground truth—conditioned on past observations—converges to a pre-specified target when the conditional quantiles of the errors are a linear function of past data. Complementary to our theory, our experiments spanning over $15$ datasets suggest that the performance improvement of our methods over baselines grows with the magnitude of the data’s dependence, even when baselines are tuned on the test set. We put these findings to the test by pre-registering an experiment for electricity demand forecasting in Texas, where our algorithms achieve over a $10$% reduction in confidence set sizes, a more than a $30$% improvement in quantile and absolute losses with respect to the observed errors, and significant outcomes on all $78$ out of $78$ pre-registered hypotheses. We provide documentation for the pypi package implementing our algorithms here: https://conformalopt.readthedocs.io/.
Felipe Areces, Christopher Mohri, Tatsunori B. Hashimoto, John C. Duchi
ICML2
2024 Learning to Reject with a Fixed Predictor: Application to Decontextualization
abstract
We study the problem of classification with a reject option for a fixed predictor, crucial to natural language processing. We introduce a new problem formulation for this scenario, and an algorithm minimizing a new surrogate loss function. We provide a complete theoretical analysis of the surrogate loss function with a strong $H$-consistency guarantee. For evaluation, we choose the \textit{decontextualization} task, and provide a manually-labelled dataset of $2\mathord,000$ examples. Our algorithm significantly outperforms the baselines considered, with a $\sim 25$% improvement in coverage when halving the error rate, which is only $\sim 3$% away from the theoretical limit.
Christopher Mohri, Daniel Andor, Eunsol Choi, Michael Collins 0001, Anqi Mao, Yutao Zhong 0002
ICLR1
2024 Language Models with Conformal Factuality Guarantees
abstract
Guaranteeing the correctness and factuality of language model (LM) outputs is a major open problem. In this work, we propose conformal factuality, a framework that can ensure high probability correctness guarantees for LMs by connecting language modeling and conformal prediction. Our insight is that the correctness of an LM output is equivalent to an uncertainty quantification problem, where the uncertainty sets are defined as the entailment set of an LM’s output. Using this connection, we show that conformal prediction in language models corresponds to a back-off algorithm that provides high probability correctness guarantees by progressively making LM outputs less specific (and expanding the associated uncertainty sets). This approach applies to any black-box LM and requires very few human-annotated samples. Evaluations of our approach on closed book QA (FActScore, NaturalQuestions) and reasoning tasks (MATH) show that our approach can provide 80-90% correctness guarantees while retaining the majority of the LM’s original output.
Christopher Mohri, Tatsunori B. Hashimoto
ICML1
2024 Cardinality-Aware Set Prediction and Top-$k$ Classification
abstract
We present a detailed study of cardinality-aware top-$k$ classification, a novel approach that aims to learn an accurate top-$k$ set predictor while maintaining a low cardinality. We introduce a new target loss function tailored to this setting that accounts for both the classification error and the cardinality of the set predicted. To optimize this loss function, we propose two families of surrogate losses: cost-sensitive comp-sum losses and cost-sensitive constrained losses. Minimizing these loss functions leads to new cardinality-aware algorithms that we describe in detail in the case of both top-$k$ and threshold-based classifiers. We establish $H$-consistency bounds for our cardinality-aware surrogate loss functions, thereby providing a strong theoretical foundation for our algorithms. We report the results of extensive experiments on CIFAR-10, CIFAR-100, ImageNet, and SVHN datasets demonstrating the effectiveness and benefits of our cardinality-aware algorithms.
Corinna Cortes, Anqi Mao, Christopher Mohri, Mehryar Mohri, Yutao Zhong 0002
NeurIPS3
2023 Theory and Algorithm for Batch Distribution Drift Problems
abstract
We study a problem of batch distribution drift motivated by several applications, which consists of determining an accurate predictor for a target time segment, for which a moderate amount of labeled samples are at one’s disposal, while leveraging past segments for which substantially more labeled samples are available. We give new algorithms for this problem guided by a new theoretical analysis and generalization bounds derived for this scenario. We further extend our results to the case where few or no labeled data is available for the period of interest. Finally, we report the results of extensive experiments demonstrating the benefits of our drifting algorithm, including comparisons with natural baselines. A by-product of our study is a principled solution to the problem of multiple-source adaptation with labeled source data and a moderate amount of target labeled data, which we briefly discuss and compare with.
Pranjal Awasthi, Corinna Cortes, Christopher Mohri
AISTATS3
2023 Two-Stage Learning to Defer with Multiple Experts
abstract
We study a two-stage scenario for learning to defer with multiple experts, which is crucial in practice for many applications. In this scenario, a predictor is derived in a first stage by training with a common loss function such as cross-entropy. In the second stage, a deferral function is learned to assign the most suitable expert to each input. We design a new family of surrogate loss functions for this scenario both in the score-based and the predictor-rejector settings and prove that they are supported by $H$-consistency bounds, which implies their Bayes-consistency. Moreover, we show that, for a constant cost function, our two-stage surrogate losses are realizable $H$-consistent. While the main focus of this work is a theoretical analysis, we also report the results of several experiments on CIFAR-10 and SVHN datasets.
Anqi Mao, Christopher Mohri, Mehryar Mohri, Yutao Zhong 0002
NeurIPS2