EDBT 2026 Demo / reviewers in the wild / expert
Kevin Du
dblp:323/9669
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2026
0009-0007-7373-3458ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Trustworthy machine learning · 48% Language models and text generation · 19% Reinforcement learning · 12% | |
| Network and information security
1 paper |
Network security · 50% Privacy and data protection · 50% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 22 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
interpretability |
2.4 | 3 | 2025 | Controllable Context Sensitivity and the Knob Behind It · ICLR 2025 How Persuasive Is Your Context? · EMNLP 2025 Generalizing Backpropagation for Gradient-Based Interpretability · ACL (1) 2023 |
Machine learning › Trustworthy machine learning
robustness |
1.8 | 2 | 2026 | It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief · ACL (1) 2026 Context versus Prior Knowledge in Language Models · ACL (1) 2024 |
Natural language and speech › Language models and text generation › prompting
prompt engineering |
1.0 | 1 | 2026 | It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief · ACL (1) 2026 |
Machine learning › Trustworthy machine learning › uncertainty estimation
bayesian uncertainty quantification |
0.9 | 1 | 2025 | Bernstein-von Mises for Adaptively Collected Data · NeurIPS 2025 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian asymptotics
bernstein-von mises theorem |
0.9 | 1 | 2025 | Bernstein-von Mises for Adaptively Collected Data · NeurIPS 2025 |
Natural language and speech › Language models and text generation
in-context learning |
0.9 | 1 | 2025 | How Persuasive Is Your Context? · EMNLP 2025 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.9 | 1 | 2025 | Bernstein-von Mises for Adaptively Collected Data · NeurIPS 2025 |
Computer vision › Vision and language › vision-language model
vision-language model evaluation |
0.9 | 1 | 2025 | Taxonomy-Aware Evaluation of Vision-Language Models · CVPR 2025 |
Computer vision › Segmentation and scene understanding › semantic segmentation
context aggregation |
0.8 | 1 | 2024 | Context versus Prior Knowledge in Language Models · ACL (1) 2024 |
Machine learning › Reinforcement learning › deep reinforcement learning
alphazero |
0.7 | 1 | 2023 | AlphaSnake: Policy Iteration on a Nondeterministic NP-Hard Markov Decision Process (Student Abstract) · AAAI 2023 |
Machine learning › Trustworthy machine learning › interpretability › attribution methods
feature attribution |
0.7 | 1 | 2023 | Generalizing Backpropagation for Gradient-Based Interpretability · ACL (1) 2023 |
Machine learning › Reinforcement learning
markov decision process |
0.7 | 1 | 2023 | AlphaSnake: Policy Iteration on a Nondeterministic NP-Hard Markov Decision Process (Student Abstract) · AAAI 2023 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search |
0.7 | 1 | 2023 | AlphaSnake: Policy Iteration on a Nondeterministic NP-Hard Markov Decision Process (Student Abstract) · AAAI 2023 |
Information retrieval › document retrieval › domain-specific retrieval › biomedical information retrieval
health search |
0.6 | 1 | 2022 | Inconsistent Ranking Assumptions in Medical Search and Their Downstream Consequences · SIGIR 2022 |
Information retrieval › ranking
ranking calibration |
0.6 | 1 | 2022 | Inconsistent Ranking Assumptions in Medical Search and Their Downstream Consequences · SIGIR 2022 |
Machine learning › Reinforcement learning
adaptive data collection |
0.3 | 1 | 2025 | Bernstein-von Mises for Adaptively Collected Data · NeurIPS 2025 |
Machine learning › Reinforcement learning
bandit |
0.3 | 1 | 2025 | Bernstein-von Mises for Adaptively Collected Data · NeurIPS 2025 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.3 | 1 | 2025 | Controllable Context Sensitivity and the Knob Behind It · ICLR 2025 |
Machine learning › Deep learning architectures and training
backpropagation |
0.2 | 1 | 2023 | Generalizing Backpropagation for Gradient-Based Interpretability · ACL (1) 2023 |
Graph algorithms and graph theory › graph theory › hamiltonicity
hamiltonian cycle |
0.2 | 1 | 2023 | AlphaSnake: Policy Iteration on a Nondeterministic NP-Hard Markov Decision Process (Student Abstract) · AAAI 2023 |
Information retrieval › retrieval models
neural retrieval |
0.2 | 1 | 2022 | Inconsistent Ranking Assumptions in Medical Search and Their Downstream Consequences · SIGIR 2022 |
Information retrieval › ranking
relevance estimation |
0.2 | 1 | 2022 | Inconsistent Ranking Assumptions in Medical Search and Their Downstream Consequences · SIGIR 2022 |
Methods — techniques the papers use, named apart from their topics
certificate analysis · 1.5TLS connection log analysis · 1.5statistical analysis · 1.0benchmark construction · 1.0wasserstein distance · 0.9text similarity measures · 0.9taxonomy mapping · 0.9targeted persuasion score · 0.9linear time layer selection algorithm · 0.9fine-tuning · 0.9distributional analysis · 0.9asymptotic analysis · 0.9policy iteration · 0.7monte carlo tree search · 0.7re-ranking · 0.6none-of-the-above classification · 0.6cutoff prediction · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of BeliefabstractUsers frequently express their beliefs to large language models (LLMs).In some situations, it is ideal for the LLM to accept this contextual information as true, while in others, it is ideal to stick to prior knowledge.Users' expressions of belief (EoBs) can take linguistically diverse forms-using presuppositions, evidential and certainty markers, or varied toneseach of which may have a different persuasiveness over the LLMs.We introduce a benchmark to systematically evaluate how different EoBs affect whether models follow context versus prior knowledge.We propose a typology grounded in four linguistically motivated dimensions: form, evidentiality, epistemic stance, and tone, spanning 19 fine-grained types.By pairing these EoBs with world knowledge facts, we generate controlled EoB-query pairs that isolate the effect of linguistic variation.We use our benchmark to evaluate 18 LLMs that differ in architecture (Llama3, Qwen3, Gemma3), scale (1B-30B parameters), and training stages (base vs instruct).We identify meaningful variations in response behavior across these axes: For example, bigger models and instruction models tend to be less context-following than smaller models and base models.We further identify specific EoBs that statistically significantly persuade LMs more consistently than others.These systematic patterns in how linguistic framing affects LLM context integration serve to evaluate model robustness and inform best practices for prompt engineering.We publicly release code and data used in this project. Kevin Du, Clara Kümpel, Michelle Wastl, Alex Warstadt |
ACL (1) | 1 |
| 2025 | Taxonomy-Aware Evaluation of Vision-Language ModelsabstractWhen a vision–language model (VLM) is prompted to identify an entity depicted in an image, it may answer "I see a conifer," rather than the specific label NORWAY SPRUCE. This raises two issues for evaluation: Firstly, the unconstrained generated text needs to be mapped to the evaluation label space (i.e., CONIFER). Secondly, a useful classification measure should give partial credit to lessspecific, but not incorrect, answers (NORWAY SPRUCE being a type of CONIFER). To meet these requirements, we propose a framework for evaluating unconstrained text predictions such as those generated from a vision–language model against a taxonomy. Specifically, we propose the use of hierarchical precision and recall measures to assess the level of correctness and specificity of predictions with regard to a taxonomy. Experimentally, we first show that existing text similarity measures do not capture taxonomic similarity well. We then develop and compare different methods to map textual VLM predictions onto a taxonomy. This allows us to compute hierarchical similarity measures between the generated text and the ground truth labels. Finally, we analyze modern VLMs on fine-grained visual classification tasks based on our proposed taxonomic evaluation scheme. Data and code are made available at https://github.com/vesteinn/vlm-eval. Vésteinn Snæbjarnarson, Kevin Du, Niklas Stoehr, Serge J. Belongie, Ryan Cotterell, Nico Lang, Stella Frank |
CVPR | 2 |
| 2025 | How Persuasive Is Your Context?abstractTwo central capabilities of language models (LMs) are: (i) drawing on prior knowledge about entities, which allows them to answer queries such as What's the official language of Austria?, and (ii) adapting to new information provided in context, e.g., Pretend the official language of Austria is Tagalog., that is pre-pended to the question.In this article, we introduce targeted persuasion score (TPS), designed to quantify how persuasive a given context is to an LM where persuasion is operationalized as the ability of the context to alter the LM's answer to the question.In contrast to evaluating persuasiveness only by inspecting the greedily decoded answer under the model, TPS provides a more fine-grained view of model behavior.Based on the Wasserstein distance, TPS measures how much a context shifts a model's original answer distribution toward a target distribution.Empirically, through a series of experiments, we show that TPS captures a more nuanced notion of persuasiveness than previously proposed metrics. Kevin Du, Alexander Miserlis Hoyle, Ryan Cotterell |
EMNLP | 2 |
| 2025 | Controllable Context Sensitivity and the Knob Behind ItabstractWhen making predictions, a language model must trade off how much it relies on its context vs. its prior knowledge.
Choosing how sensitive the model is to its context is a fundamental functionality, as it enables the model to excel at tasks like retrieval-augmented generation and question-answering.
In this paper, we search for a knob which controls this sensitivity, determining whether language models answer from the context or their prior knowledge.
To guide this search, we design a task for controllable context sensitivity.
In this task, we first feed the model a context ("Paris is in England") and a question ("Where is Paris?"); we then instruct the model to either use its prior or contextual knowledge and evaluate whether it generates the correct answer for both intents (either "France" or "England").
When fine-tuned on this task, instruct versions of Llama-3.1, Mistral-v0.3, and Gemma-2 can solve it with high accuracy (85-95%).
Analyzing these high-performing models, we narrow down which layers may be important to context sensitivity using a novel linear time algorithm.
Then, in each model, we identify a 1-D subspace in a single layer that encodes whether the model follows context or prior knowledge.
Interestingly, while we identify this subspace in a fine-tuned model, we find that the exact same subspace serves as an effective knob in not only that model but also non-fine-tuned instruct and base models of that model family.
Finally, we show a strong correlation between a model's performance and how distinctly it separates context-agreeing from context-ignoring answers in this subspace.
These results suggest a single fundamental subspace facilitates how the model chooses between context and prior knowledge. Julian Minder, Kevin Du, Niklas Stoehr, Giovanni Monea, Chris Wendler, Robert West 0001, Ryan Cotterell |
ICLR | 2 |
| 2025 | Bernstein-von Mises for Adaptively Collected DataabstractUncertainty quantification (UQ) for adaptively collected data, such as that coming from adaptive experiments, bandits, or reinforcement learning, is necessary for critical elements of data collection such as ensuring safety and conducting after-study inference. The data's adaptivity creates significant challenges for frequentist UQ, yet Bayesian UQ remains the same as if the data were independent and identically distributed (i.i.d.), making it an appealing and commonly used approach. Bayesian UQ requires the (correct) specification of a prior distribution while frequentist UQ does not, but for i.i.d. data the celebrated Bernstein–von Mises theorem shows that as the sample size grows, the prior `washes out' and Bayesian UQ becomes frequentist-valid, implying that the choice of prior need not be a major impediment to Bayesian UQ as it makes no difference asymptotically. This paper for the first time extends the Bernstein–von Mises theorem to adaptively collected data, proving asymptotic equivalence between Bayesian UQ and Wald-type frequentist UQ in this challenging setting. Our results do not require the standard stability condition for validity of Wald-type frequentist UQ, and thus provide positive results on frequentist validity of Bayesian UQ under stability. Counterintuitively however, they also provide a negative result that Bayesian UQ is not asymptotically frequentist valid when stability fails, despite the fact that the prior washes out and Bayesian UQ asymptotically matches standard Wald-type frequentist UQ. We empirically validate our theory (positive and negative) via a range of simulations. Kevin Du, Yash Nair, Lucas Janson |
NeurIPS | 1 |
| 2025 | Design of a Contingent Decision Equalizer Using Analog-in-Time Processing for PAMN Wireline ReceiversabstractThe need to rapidly increase per-lane data rates to meet system bandwidth demands has led to the adoption of digital equalization schemes in wireline transceivers. However, for short reach applications, mixed-mode designs offer better power and area efficiency, especially if their complexity explosion with data rates can be alleviated. This paper presents a new hybrid linear/nonlinear equalizer that unfolds the popular recurrent equalization architecture to remove any fundamental timing constraints in the wireline receiver. The modular multi-stage architecture drastically reduces the receiver complexity and enable real-time architecture tuning to improve energy efficiency. This paves the way for continual data rate enhancement with improved circuit performance. To ensure reliable performance in advanced CMOS processes, time-domain analog signal processing is used. This choice also enables linear energy scaling with data rate and allows proportional power scaling with degree of equalization. A 3-tap realization of the receiver fabricated in 28 nm CMOS process achieved a data rate of 52 Gb/s with an efficiency of 1.0 pJ/b and BER below 1e-4 with moderate ISI where channel loss at Nyquist is less than 12 dB. Kevin Du, Mohamed O. Abouzeid, Tawfiq Musah |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | Context versus Prior Knowledge in Language ModelsabstractKevin Du, Vésteinn Snæbjarnarson, Niklas Stoehr, Jennifer White, Aaron Schein, Ryan Cotterell. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Kevin Du, Vésteinn Snæbjarnarson, Niklas Stoehr, Jennifer C. White, Aaron Schein, Ryan Cotterell |
ACL (1) | 1 |
| 2024 | Mutual TLS in Practice: A Deep Dive into Certificate Configurations and Privacy IssuesabstractTransport Layer Security (TLS) is widely recognized as the essential protocol for securing Internet communications. While numerous studies have focused on investigating server certificates used in TLS connections, our study delves into the less explored territory of mutual TLS (mTLS) where both parties need to provide certificates to each other. By utilizing TLS connection logs collected from a large campus network over 23 months, we identify over 2.2 million unique server certificates and over 3.4 million unique client certificates used in over 1.2 billion mutual TLS connections. By jointly analyzing TLS connection data (e.g., port numbers) and certificate data (e.g., issuers for server/client certificates), we quantify the prevalent use of untrusted certificates and uncover potential security concerns resulting from misconfigured certificates, sharing of certificates between servers and clients, and long-expired certificates. Furthermore, we present the first in-depth study on the wide range of information included in CommonName (CN) and Subject Alternative Name (SAN), drawing comparison between client and server certificates, as well as revealing sensitive information. Hongying Dong, Yizhe Zhang 0006, Hyeonmin Lee, Kevin Du, Guancheng Tu, Yixin Sun 0004 |
IMC | 4 |
| 2023 | AlphaSnake: Policy Iteration on a Nondeterministic NP-Hard Markov Decision Process (Student Abstract)abstractReinforcement learning has been used to approach well-known NP-hard combinatorial problems in graph theory. Among these, Hamiltonian cycle problems are exceptionally difficult to analyze, even when restricted to individual instances of structurally complex graphs. In this paper, we use Monte Carlo Tree Search (MCTS), the search algorithm behind many state-of-the-art reinforcement learning algorithms such as AlphaZero, to create autonomous agents that learn to play the game of Snake, a game centered on properties of Hamiltonian cycles on grid graphs. The game of Snake can be formulated as a single-player discounted Markov Decision Process (MDP), where the agent must behave optimally in a stochastic environment. Determining the optimal policy for Snake, defined as the policy that maximizes the probability of winning -- or win rate -- with higher priority and minimizes the expected number of time steps to win with lower priority, is conjectured to be NP-hard. Performance-wise, compared to prior work in the Snake game, our algorithm is the first to achieve a win rate over 0.5 (a uniform random policy achieves a win rate < 2.57 x 10^{-15}), demonstrating the versatility of AlphaZero in tackling NP-hard problems. Kevin Du, Ian Gemp, Yi Wu 0013 |
AAAI | 1 |
| 2023 | Generalizing Backpropagation for Gradient-Based InterpretabilityabstractMany popular feature-attribution methods for interpreting deep neural networks rely on computing the gradients of a model's output with respect to its inputs.While these methods can indicate which input features may be important for the model's prediction, they reveal little about the inner workings of the model itself.In this paper, we observe that the gradient computation of a model is a special case of a more general formulation using semirings.This observation allows us to generalize the backpropagation algorithm to efficiently compute other interpretable statistics about the gradient graph of a neural network, such as the highest-weighted path and entropy.We implement this generalized algorithm, evaluate it on synthetic datasets to better understand the statistics it computes, and apply it to study BERT's behavior on the subject-verb number agreement task (SVA).With this method, we (a) validate that the amount of gradient flow through a component of a model reflects its importance to a prediction and (b) for SVA, identify which pathways of the self-attention mechanism are most important. Kevin Du, Lucas Torroba Hennigen, Niklas Stoehr, Alex Warstadt, Ryan Cotterell |
ACL (1) | 1 |
| 2022 | Inconsistent Ranking Assumptions in Medical Search and Their Downstream ConsequencesabstractGiven a query, neural retrieval models predict point estimates of relevance for each document; however, a significant drawback of relying solely on point estimates is that they contain no indication of the model's confidence in its predictions. Despite this lack of information, downstream methods such as reranking, cutoff prediction, and none-of-the-above classification are still able to learn effective functions to accomplish their respective tasks. Unfortunately, these downstream methods can suffer poor performance when the initial ranking model loses confidence in its score predictions. This becomes increasingly important in high-stakes settings, such as medical searches that can influence health decision making. Kevin Du, Bhaskar Mitra 0001, Laura Mercurio, Navid Rekabsaz, Carsten Eickhoff |
SIGIR | 2 |