Kevin Du

dblp:323/9669 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2026
0009-0007-7373-3458ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Trustworthy machine learning · 48% Language models and text generation · 19% Reinforcement learning · 12%
Network and information security
1 paper
Network security · 50% Privacy and data protection · 50%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 22 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
2.432025
Controllable Context Sensitivity and the Knob Behind It · ICLR 2025
How Persuasive Is Your Context? · EMNLP 2025
Generalizing Backpropagation for Gradient-Based Interpretability · ACL (1) 2023
Machine learning › Trustworthy machine learning
robustness
1.822026
It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief · ACL (1) 2026
Context versus Prior Knowledge in Language Models · ACL (1) 2024
Natural language and speech › Language models and text generation › prompting
prompt engineering
1.012026
It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief · ACL (1) 2026
Machine learning › Trustworthy machine learning › uncertainty estimation
bayesian uncertainty quantification
0.912025
Bernstein-von Mises for Adaptively Collected Data · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian asymptotics
bernstein-von mises theorem
0.912025
Bernstein-von Mises for Adaptively Collected Data · NeurIPS 2025
Natural language and speech › Language models and text generation
in-context learning
0.912025
How Persuasive Is Your Context? · EMNLP 2025
Machine learning › Trustworthy machine learning
uncertainty estimation
0.912025
Bernstein-von Mises for Adaptively Collected Data · NeurIPS 2025
Computer vision › Vision and language › vision-language model
vision-language model evaluation
0.912025
Taxonomy-Aware Evaluation of Vision-Language Models · CVPR 2025
Computer vision › Segmentation and scene understanding › semantic segmentation
context aggregation
0.812024
Context versus Prior Knowledge in Language Models · ACL (1) 2024
Machine learning › Reinforcement learning › deep reinforcement learning
alphazero
0.712023
AlphaSnake: Policy Iteration on a Nondeterministic NP-Hard Markov Decision Process (Student Abstract) · AAAI 2023
Machine learning › Trustworthy machine learning › interpretability › attribution methods
feature attribution
0.712023
Generalizing Backpropagation for Gradient-Based Interpretability · ACL (1) 2023
Machine learning › Reinforcement learning
markov decision process
0.712023
AlphaSnake: Policy Iteration on a Nondeterministic NP-Hard Markov Decision Process (Student Abstract) · AAAI 2023
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search
0.712023
AlphaSnake: Policy Iteration on a Nondeterministic NP-Hard Markov Decision Process (Student Abstract) · AAAI 2023
Information retrieval › document retrieval › domain-specific retrieval › biomedical information retrieval
health search
0.612022
Inconsistent Ranking Assumptions in Medical Search and Their Downstream Consequences · SIGIR 2022
Information retrieval › ranking
ranking calibration
0.612022
Inconsistent Ranking Assumptions in Medical Search and Their Downstream Consequences · SIGIR 2022
Machine learning › Reinforcement learning
adaptive data collection
0.312025
Bernstein-von Mises for Adaptively Collected Data · NeurIPS 2025
Machine learning › Reinforcement learning
bandit
0.312025
Bernstein-von Mises for Adaptively Collected Data · NeurIPS 2025
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.312025
Controllable Context Sensitivity and the Knob Behind It · ICLR 2025
Machine learning › Deep learning architectures and training
backpropagation
0.212023
Generalizing Backpropagation for Gradient-Based Interpretability · ACL (1) 2023
Graph algorithms and graph theory › graph theory › hamiltonicity
hamiltonian cycle
0.212023
AlphaSnake: Policy Iteration on a Nondeterministic NP-Hard Markov Decision Process (Student Abstract) · AAAI 2023
Information retrieval › retrieval models
neural retrieval
0.212022
Inconsistent Ranking Assumptions in Medical Search and Their Downstream Consequences · SIGIR 2022
Information retrieval › ranking
relevance estimation
0.212022
Inconsistent Ranking Assumptions in Medical Search and Their Downstream Consequences · SIGIR 2022

Methods — techniques the papers use, named apart from their topics

certificate analysis · 1.5TLS connection log analysis · 1.5statistical analysis · 1.0benchmark construction · 1.0wasserstein distance · 0.9text similarity measures · 0.9taxonomy mapping · 0.9targeted persuasion score · 0.9linear time layer selection algorithm · 0.9fine-tuning · 0.9distributional analysis · 0.9asymptotic analysis · 0.9policy iteration · 0.7monte carlo tree search · 0.7re-ranking · 0.6none-of-the-above classification · 0.6cutoff prediction · 0.6
YearPublicationVenuePosition
2026 It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief
abstract
Users frequently express their beliefs to large language models (LLMs).In some situations, it is ideal for the LLM to accept this contextual information as true, while in others, it is ideal to stick to prior knowledge.Users' expressions of belief (EoBs) can take linguistically diverse forms-using presuppositions, evidential and certainty markers, or varied toneseach of which may have a different persuasiveness over the LLMs.We introduce a benchmark to systematically evaluate how different EoBs affect whether models follow context versus prior knowledge.We propose a typology grounded in four linguistically motivated dimensions: form, evidentiality, epistemic stance, and tone, spanning 19 fine-grained types.By pairing these EoBs with world knowledge facts, we generate controlled EoB-query pairs that isolate the effect of linguistic variation.We use our benchmark to evaluate 18 LLMs that differ in architecture (Llama3, Qwen3, Gemma3), scale (1B-30B parameters), and training stages (base vs instruct).We identify meaningful variations in response behavior across these axes: For example, bigger models and instruction models tend to be less context-following than smaller models and base models.We further identify specific EoBs that statistically significantly persuade LMs more consistently than others.These systematic patterns in how linguistic framing affects LLM context integration serve to evaluate model robustness and inform best practices for prompt engineering.We publicly release code and data used in this project.
Kevin Du, Clara Kümpel, Michelle Wastl, Alex Warstadt
ACL (1)1
2025 Taxonomy-Aware Evaluation of Vision-Language Models
abstract
When a vision–language model (VLM) is prompted to identify an entity depicted in an image, it may answer "I see a conifer," rather than the specific label NORWAY SPRUCE. This raises two issues for evaluation: Firstly, the unconstrained generated text needs to be mapped to the evaluation label space (i.e., CONIFER). Secondly, a useful classification measure should give partial credit to lessspecific, but not incorrect, answers (NORWAY SPRUCE being a type of CONIFER). To meet these requirements, we propose a framework for evaluating unconstrained text predictions such as those generated from a vision–language model against a taxonomy. Specifically, we propose the use of hierarchical precision and recall measures to assess the level of correctness and specificity of predictions with regard to a taxonomy. Experimentally, we first show that existing text similarity measures do not capture taxonomic similarity well. We then develop and compare different methods to map textual VLM predictions onto a taxonomy. This allows us to compute hierarchical similarity measures between the generated text and the ground truth labels. Finally, we analyze modern VLMs on fine-grained visual classification tasks based on our proposed taxonomic evaluation scheme. Data and code are made available at https://github.com/vesteinn/vlm-eval.
Vésteinn Snæbjarnarson, Kevin Du, Niklas Stoehr, Serge J. Belongie, Ryan Cotterell, Nico Lang, Stella Frank
CVPR2
2025 How Persuasive Is Your Context?
abstract
Two central capabilities of language models (LMs) are: (i) drawing on prior knowledge about entities, which allows them to answer queries such as What's the official language of Austria?, and (ii) adapting to new information provided in context, e.g., Pretend the official language of Austria is Tagalog., that is pre-pended to the question.In this article, we introduce targeted persuasion score (TPS), designed to quantify how persuasive a given context is to an LM where persuasion is operationalized as the ability of the context to alter the LM's answer to the question.In contrast to evaluating persuasiveness only by inspecting the greedily decoded answer under the model, TPS provides a more fine-grained view of model behavior.Based on the Wasserstein distance, TPS measures how much a context shifts a model's original answer distribution toward a target distribution.Empirically, through a series of experiments, we show that TPS captures a more nuanced notion of persuasiveness than previously proposed metrics.
Kevin Du, Alexander Miserlis Hoyle, Ryan Cotterell
EMNLP2
2025 Controllable Context Sensitivity and the Knob Behind It
abstract
When making predictions, a language model must trade off how much it relies on its context vs. its prior knowledge. Choosing how sensitive the model is to its context is a fundamental functionality, as it enables the model to excel at tasks like retrieval-augmented generation and question-answering. In this paper, we search for a knob which controls this sensitivity, determining whether language models answer from the context or their prior knowledge. To guide this search, we design a task for controllable context sensitivity. In this task, we first feed the model a context ("Paris is in England") and a question ("Where is Paris?"); we then instruct the model to either use its prior or contextual knowledge and evaluate whether it generates the correct answer for both intents (either "France" or "England"). When fine-tuned on this task, instruct versions of Llama-3.1, Mistral-v0.3, and Gemma-2 can solve it with high accuracy (85-95%). Analyzing these high-performing models, we narrow down which layers may be important to context sensitivity using a novel linear time algorithm. Then, in each model, we identify a 1-D subspace in a single layer that encodes whether the model follows context or prior knowledge. Interestingly, while we identify this subspace in a fine-tuned model, we find that the exact same subspace serves as an effective knob in not only that model but also non-fine-tuned instruct and base models of that model family. Finally, we show a strong correlation between a model's performance and how distinctly it separates context-agreeing from context-ignoring answers in this subspace. These results suggest a single fundamental subspace facilitates how the model chooses between context and prior knowledge.
Julian Minder, Kevin Du, Niklas Stoehr, Giovanni Monea, Chris Wendler, Robert West 0001, Ryan Cotterell
ICLR2
2025 Bernstein-von Mises for Adaptively Collected Data
abstract
Uncertainty quantification (UQ) for adaptively collected data, such as that coming from adaptive experiments, bandits, or reinforcement learning, is necessary for critical elements of data collection such as ensuring safety and conducting after-study inference. The data's adaptivity creates significant challenges for frequentist UQ, yet Bayesian UQ remains the same as if the data were independent and identically distributed (i.i.d.), making it an appealing and commonly used approach. Bayesian UQ requires the (correct) specification of a prior distribution while frequentist UQ does not, but for i.i.d. data the celebrated Bernstein–von Mises theorem shows that as the sample size grows, the prior `washes out' and Bayesian UQ becomes frequentist-valid, implying that the choice of prior need not be a major impediment to Bayesian UQ as it makes no difference asymptotically. This paper for the first time extends the Bernstein–von Mises theorem to adaptively collected data, proving asymptotic equivalence between Bayesian UQ and Wald-type frequentist UQ in this challenging setting. Our results do not require the standard stability condition for validity of Wald-type frequentist UQ, and thus provide positive results on frequentist validity of Bayesian UQ under stability. Counterintuitively however, they also provide a negative result that Bayesian UQ is not asymptotically frequentist valid when stability fails, despite the fact that the prior washes out and Bayesian UQ asymptotically matches standard Wald-type frequentist UQ. We empirically validate our theory (positive and negative) via a range of simulations.
Kevin Du, Yash Nair, Lucas Janson
NeurIPS1
2025 Design of a Contingent Decision Equalizer Using Analog-in-Time Processing for PAMN Wireline Receivers
abstract
The need to rapidly increase per-lane data rates to meet system bandwidth demands has led to the adoption of digital equalization schemes in wireline transceivers. However, for short reach applications, mixed-mode designs offer better power and area efficiency, especially if their complexity explosion with data rates can be alleviated. This paper presents a new hybrid linear/nonlinear equalizer that unfolds the popular recurrent equalization architecture to remove any fundamental timing constraints in the wireline receiver. The modular multi-stage architecture drastically reduces the receiver complexity and enable real-time architecture tuning to improve energy efficiency. This paves the way for continual data rate enhancement with improved circuit performance. To ensure reliable performance in advanced CMOS processes, time-domain analog signal processing is used. This choice also enables linear energy scaling with data rate and allows proportional power scaling with degree of equalization. A 3-tap realization of the receiver fabricated in 28 nm CMOS process achieved a data rate of 52 Gb/s with an efficiency of 1.0 pJ/b and BER below 1e-4 with moderate ISI where channel loss at Nyquist is less than 12 dB.
Kevin Du, Mohamed O. Abouzeid, Tawfiq Musah
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 Context versus Prior Knowledge in Language Models
abstract
Kevin Du, Vésteinn Snæbjarnarson, Niklas Stoehr, Jennifer White, Aaron Schein, Ryan Cotterell. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Kevin Du, Vésteinn Snæbjarnarson, Niklas Stoehr, Jennifer C. White, Aaron Schein, Ryan Cotterell
ACL (1)1
2024 Mutual TLS in Practice: A Deep Dive into Certificate Configurations and Privacy Issues
abstract
Transport Layer Security (TLS) is widely recognized as the essential protocol for securing Internet communications. While numerous studies have focused on investigating server certificates used in TLS connections, our study delves into the less explored territory of mutual TLS (mTLS) where both parties need to provide certificates to each other. By utilizing TLS connection logs collected from a large campus network over 23 months, we identify over 2.2 million unique server certificates and over 3.4 million unique client certificates used in over 1.2 billion mutual TLS connections. By jointly analyzing TLS connection data (e.g., port numbers) and certificate data (e.g., issuers for server/client certificates), we quantify the prevalent use of untrusted certificates and uncover potential security concerns resulting from misconfigured certificates, sharing of certificates between servers and clients, and long-expired certificates. Furthermore, we present the first in-depth study on the wide range of information included in CommonName (CN) and Subject Alternative Name (SAN), drawing comparison between client and server certificates, as well as revealing sensitive information.
Hongying Dong, Yizhe Zhang 0006, Hyeonmin Lee, Kevin Du, Guancheng Tu, Yixin Sun 0004
IMC4
2023 AlphaSnake: Policy Iteration on a Nondeterministic NP-Hard Markov Decision Process (Student Abstract)
abstract
Reinforcement learning has been used to approach well-known NP-hard combinatorial problems in graph theory. Among these, Hamiltonian cycle problems are exceptionally difficult to analyze, even when restricted to individual instances of structurally complex graphs. In this paper, we use Monte Carlo Tree Search (MCTS), the search algorithm behind many state-of-the-art reinforcement learning algorithms such as AlphaZero, to create autonomous agents that learn to play the game of Snake, a game centered on properties of Hamiltonian cycles on grid graphs. The game of Snake can be formulated as a single-player discounted Markov Decision Process (MDP), where the agent must behave optimally in a stochastic environment. Determining the optimal policy for Snake, defined as the policy that maximizes the probability of winning -- or win rate -- with higher priority and minimizes the expected number of time steps to win with lower priority, is conjectured to be NP-hard. Performance-wise, compared to prior work in the Snake game, our algorithm is the first to achieve a win rate over 0.5 (a uniform random policy achieves a win rate < 2.57 x 10^{-15}), demonstrating the versatility of AlphaZero in tackling NP-hard problems.
Kevin Du, Ian Gemp, Yi Wu 0013
AAAI1
2023 Generalizing Backpropagation for Gradient-Based Interpretability
abstract
Many popular feature-attribution methods for interpreting deep neural networks rely on computing the gradients of a model's output with respect to its inputs.While these methods can indicate which input features may be important for the model's prediction, they reveal little about the inner workings of the model itself.In this paper, we observe that the gradient computation of a model is a special case of a more general formulation using semirings.This observation allows us to generalize the backpropagation algorithm to efficiently compute other interpretable statistics about the gradient graph of a neural network, such as the highest-weighted path and entropy.We implement this generalized algorithm, evaluate it on synthetic datasets to better understand the statistics it computes, and apply it to study BERT's behavior on the subject-verb number agreement task (SVA).With this method, we (a) validate that the amount of gradient flow through a component of a model reflects its importance to a prediction and (b) for SVA, identify which pathways of the self-attention mechanism are most important.
Kevin Du, Lucas Torroba Hennigen, Niklas Stoehr, Alex Warstadt, Ryan Cotterell
ACL (1)1
2022 Inconsistent Ranking Assumptions in Medical Search and Their Downstream Consequences
abstract
Given a query, neural retrieval models predict point estimates of relevance for each document; however, a significant drawback of relying solely on point estimates is that they contain no indication of the model's confidence in its predictions. Despite this lack of information, downstream methods such as reranking, cutoff prediction, and none-of-the-above classification are still able to learn effective functions to accomplish their respective tasks. Unfortunately, these downstream methods can suffer poor performance when the initial ranking model loses confidence in its score predictions. This becomes increasingly important in high-stakes settings, such as medical searches that can influence health decision making.
Kevin Du, Bhaskar Mitra 0001, Laura Mercurio, Navid Rekabsaz, Carsten Eickhoff
SIGIR2