Trenton Chang

dblp:322/4029 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0003-4679-5841ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Trustworthy machine learning · 73% Language models and text generation · 24% Probabilistic and Bayesian machine learning · 3%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational social science and digital humanities · 54% Medical and health informatics · 46%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
1.622025
Disentangling misreporting from genuine adaptation in strategic settings: a causal approach · NeurIPS 2025
From Biased Selective Labels to Pseudo-Labels: An Expectation-Maximization Framework for Learning from Biased Decisions · ICML 2024
Machine learning › Trustworthy machine learning
strategic behavior
1.622025
Disentangling misreporting from genuine adaptation in strategic settings: a causal approach · NeurIPS 2025
Who's Gaming the System? A Causally-Motivated Approach for Detecting Strategic Adaptation · NeurIPS 2024
Natural language and speech › Language models and text generation
alignment
1.012026
A Course Correction in Steerability Evaluation: Revealing Miscalibration and Side Effects in LLMs · AAAI 2026
Machine learning › Trustworthy machine learning
interpretability
1.012026
A Course Correction in Steerability Evaluation: Revealing Miscalibration and Side Effects in LLMs · AAAI 2026
Natural language and speech › Language models and text generation
large language model
1.012026
A Course Correction in Steerability Evaluation: Revealing Miscalibration and Side Effects in LLMs · AAAI 2026
Machine learning › Trustworthy machine learning › fairness
causal fairness
0.912025
Disentangling misreporting from genuine adaptation in strategic settings: a causal approach · NeurIPS 2025
Machine learning › Trustworthy machine learning › performative prediction
strategic classification
0.912025
Disentangling misreporting from genuine adaptation in strategic settings: a causal approach · NeurIPS 2025
Computational social science and digital humanities
algorithmic decision-making
0.912025
Disentangling misreporting from genuine adaptation in strategic settings: a causal approach · NeurIPS 2025
Medical and health informatics › clinical informatics › clinical AI
clinical machine learning
0.812024
From Biased Selective Labels to Pseudo-Labels: An Expectation-Maximization Framework for Learning from Biased Decisions · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
expectation-maximization
0.212024
From Biased Selective Labels to Pseudo-Labels: An Expectation-Maximization Framework for Learning from Biased Decisions · ICML 2024

Methods — techniques the papers use, named apart from their topics

identifiability analysis · 2.5causal inference · 1.7expectation-maximization · 1.5causal models · 1.5reinforcement learning fine-tuning · 1.0prompt engineering · 1.0best-of-n sampling · 1.0causal effect estimation · 0.8
YearPublicationVenuePosition
2026 A Course Correction in Steerability Evaluation: Revealing Miscalibration and Side Effects in LLMs
abstract
Despite advances in large language models (LLMs) on reasoning and instruction-following benchmarks, it is unclear whether they can reliably produce outputs aligned with a variety of user goals, a concept called steerability. We highlight two gaps in current LLM evaluations for assessing steerability. First, many benchmarks are built with past LLM chats and text scraped from the Internet, which may skew towards common requests, underrepresenting less-common requests by potential users. Second, prior work measures performance as a scalar, which could conceal behavioral shifts in LLM outputs in open-ended generation. To mitigate these gaps, we introduce a framework based on a multi-dimensional goal space that models user goals and LLM outputs as vectors with dimensions corresponding to text attributes (e.g., reading difficulty). Applied to a text-rewriting task, we find that current LLMs induce intended changes or "side-effects" to text attributes, impeding steerability. Interventions to improve steerability, such as prompt engineering, best-of-N sampling, and reinforcement learning fine-tuning, have varying effectiveness, yet side effects remain problematic. Our findings suggest that even strong LLMs struggle with steerability, and existing alignment strategies may be insufficient.
Trenton Chang, Tobias Schnabel, Adith Swaminathan, Jenna Wiens
AAAI1
2025 Disentangling misreporting from genuine adaptation in strategic settings: a causal approach
abstract
In settings where ML models are used to inform the allocation of resources, agents affected by the allocation decisions might have an incentive to strategically change their features to secure better outcomes. While prior work has studied strategic responses broadly, disentangling misreporting from genuine adaptation remains a fundamental challenge. In this paper, we propose a causally-motivated approach to identify and quantify how much an agent misreports on average by distinguishing deceptive changes in their features from genuine adaptation. Our key insight is that, unlike genuine adaptation, misreported features do not causally affect downstream variables (i.e., causal descendants). We exploit this asymmetry by comparing the causal effect of misreported features on their causal descendants as derived from manipulated datasets against those from unmanipulated datasets. We formally prove identifiability of the misreporting rate and characterize the variance of our estimator. We empirically validate our theoretical results using a semi-synthetic and real Medicare dataset with misreported data, demonstrating that our approach can be employed to identify misreporting in real-world scenarios.
Dylan Zapzalka, Trenton Chang, Lindsay A. Warrenburg, Sae-Hwan Park, Daniel K. Shenfeld, Ravi B. Parikh, Jenna Wiens, Maggie Makar
NeurIPS2
2024 From Biased Selective Labels to Pseudo-Labels: An Expectation-Maximization Framework for Learning from Biased Decisions
abstract
Selective labels occur when label observations are subject to a decision-making process; e.g., diagnoses that depend on the administration of laboratory tests. We study a clinically-inspired selective label problem called disparate censorship, where labeling biases vary across subgroups and unlabeled individuals are imputed as “negative” (i.e., no diagnostic test = no illness). Machine learning models naively trained on such labels could amplify labeling bias. Inspired by causal models of selective labels, we propose Disparate Censorship Expectation-Maximization (DCEM), an algorithm for learning in the presence of disparate censorship. We theoretically analyze how DCEM mitigates the effects of disparate censorship on model performance. We validate DCEM on synthetic data, showing that it improves bias mitigation (area between ROC curves) without sacrificing discriminative performance (AUC) compared to baselines. We achieve similar results in a sepsis classification task using clinical data.
Trenton Chang, Jenna Wiens
ICML1
2024 Who's Gaming the System? A Causally-Motivated Approach for Detecting Strategic Adaptation
abstract
In many settings, machine learning models may be used to inform decisions that impact individuals or entities who interact with the model. Such entities, or *agents,* may *game* model decisions by manipulating their inputs to the model to obtain better outcomes and maximize some utility. We consider a multi-agent setting where the goal is to identify the “worst offenders:” agents that are gaming most aggressively. However, identifying such agents is difficult without knowledge of their utility function. Thus, we introduce a framework in which each agent’s tendency to game is parameterized via a scalar. We show that this gaming parameter is only partially identifiable. By recasting the problem as a causal effect estimation problem where different agents represent different “treatments,” we prove that a ranking of all agents by their gaming parameters is identifiable. We present empirical results in a synthetic data study validating the usage of causal effect estimation for gaming detection and show in a case study of diagnosis coding behavior in the U.S. that our approach highlights features associated with gaming.
Trenton Chang, Lindsay A. Warrenburg, Sae-Hwan Park, Ravi B. Parikh, Maggie Makar, Jenna Wiens
NeurIPS1
2022 Neural Generation Meets Real People: Building a Social, Informative Open-Domain Dialogue Agent
abstract
Ethan A. Chi, Ashwin Paranjape, Abigail See, Caleb Chiam, Trenton Chang, Kathleen Kenealy, Swee Kiat Lim, Amelia Hardy, Chetanya Rastogi, Haojun Li, Alexander Iyabor, Yutong He, Hari Sowrirajan, Peng Qi, Kaushik Ram Sadagopan, Nguyet Minh Phu, Dilara Soylu, Jillian Tang, Avanika Narayan, Giovanni Campagna, Christopher Manning. Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2022.
Ethan A. Chi, Ashwin Paranjape, Abigail See, Caleb Chiam, Trenton Chang, Kathleen Kenealy, Swee Kiat Lim, Amelia F. Hardy, Chetanya Rastogi, Alexander Iyabor, Hari Sowrirajan, Peng Qi 0003, Kaushik Ram Sadagopan, Nguyet Minh Phu, Dilara Soylu, Jillian Tang, Avanika Narayan, Giovanni Campagna, Christopher D. Manning
SIGDIAL5