Bruce W. Lee

dblp:271/8029 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
6since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 44% Trustworthy machine learning · 28% Knowledge representation and reasoning · 14%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
1.012026
Inertia in Moral and Value Judgments of Large Language Models · ACL (1) 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning › commonsense reasoning
moral judgment
1.012026
Inertia in Moral and Value Judgments of Large Language Models · ACL (1) 2026
Natural language and speech › Language models and text generation › prompting
persona prompting
1.012026
Inertia in Moral and Value Judgments of Large Language Models · ACL (1) 2026
Natural language and speech › Language models and text generation › model steering › language model steering
activation steering
0.912025
Programming Refusal with Conditional Activation Steering · ICLR 2025
Natural language and speech › Language models and text generation
alignment
0.912025
Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs · NeurIPS 2025
Natural language and speech › Language models and text generation
controllable text generation
0.912025
Programming Refusal with Conditional Activation Steering · ICLR 2025
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.912025
Distillation Robustifies Unlearning · NeurIPS 2025
Machine learning › Trustworthy machine learning
machine unlearning
0.912025
Distillation Robustifies Unlearning · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model safety
refusal behavior
0.912025
Programming Refusal with Conditional Activation Steering · ICLR 2025
Machine learning › Trustworthy machine learning
robustness
0.912025
Programming Refusal with Conditional Activation Steering · ICLR 2025
Machine learning › Trustworthy machine learning › machine unlearning
robust unlearning
0.912025
Distillation Robustifies Unlearning · NeurIPS 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.812024
Language Models Don't Learn the Physical Manifestation of Language · ACL (1) 2024
Natural language and speech › Information extraction and text analysis › text classification
readability assessment
0.512021
Pushing on Text Readability Assessment: A Transformer Meets Handcrafted Linguistic Features · EMNLP (1) 2021
Machine learning › Deep learning architectures and training
transformer
0.512021
Pushing on Text Readability Assessment: A Transformer Meets Handcrafted Linguistic Features · EMNLP (1) 2021
Natural language and speech › Language models and text generation
large language model
0.312025
Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

role-play at scale · 1.0persona prompting · 1.0utility function · 0.9knowledge distillation · 0.9fine-tuning · 0.9citizen assembly · 0.9activation steering · 0.9few-shot prompting · 0.8chain-of-thought prompting · 0.8hybrid model · 0.5
YearPublicationVenuePosition
2026 Inertia in Moral and Value Judgments of Large Language Models
abstract
Large Language Models (LLMs) behave nondeterministically, and prompting has become a common method for steering their outputs.A popular strategy is to assign a persona to the model to produce more varied, contextsensitive responses, similar to how responses vary across human individuals.Against the expectation that persona prompting yields a wide range of opinions, our experiments show that LLMs keep consistent value orientations.We observe a persistent inertia in their responses, where certain moral and value dimensions (especially harm avoidance and fairness) stay skewed in one direction across persona settings.To study this, we use role-play at scale, which pairs randomized persona prompts with a macro-level analysis of model outputs.Our results point to strong internal biases and value preferences in LLMs, which we call value orientation and inertia.These models warrant scrutiny and adjustment before use in applications where balanced outputs matter.
Bruce W. Lee, Yeongheon Lee, Hyunsoo Cho
ACL (1)1
2025 Programming Refusal with Conditional Activation Steering
abstract
LLMs have shown remarkable capabilities, but precisely controlling their response behavior remains challenging. Existing activation steering methods alter LLM behavior indiscriminately, limiting their practical applicability in settings where selective responses are essential, such as content moderation or domain-specific assistants. In this paper, we propose Conditional Activation Steering (CAST), which analyzes LLM activation patterns during inference to selectively apply or withhold activation steering based on the input context. Our method is based on the observation that different categories of prompts activate distinct patterns in the model's hidden states. Using CAST, one can systematically control LLM behavior with rules like "if input is about hate speech or adult content, then refuse" or "if input is not about legal advice, then refuse." This allows for selective modification of responses to specific content while maintaining normal responses to other content, all without requiring weight optimization. We release an open-source implementation of our framework.
Bruce W. Lee, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Erik Miehling, Pierre L. Dognin, Manish Nagireddy, Amit Dhurandhar
ICLR1
2025 Distillation Robustifies Unlearning
abstract
Current LLM unlearning methods are not robust. A few steps of finetuning can revert their effects. We begin by showing that this is true even for an idealized form of unlearning: training to imitate a model that was never trained on unwanted information. This shows that training a model can drastically modify its input-output behavior while leaving its underlying capabilities intact. In light of this dynamic, we show our main result. Training a randomly initialized student on the outputs of an unlearned model transfers behaviors while leaving latent capabilities behind. In short, distillation robustifies unlearning. Based on this result, we propose Unlearn-Noise-Distill-on-Outputs (UNDO), a scalable method that distills an unlearned model into a noised copy of itself. UNDO introduces a tunable tradeoff between compute cost and robustness, establishing a new Pareto frontier on synthetic language and arithmetic tasks. At its strongest setting, UNDO matches the robustness of a model retrained from scratch with perfect data filtering while using only 60-80% of the compute and requiring only 0.01% of the pretraining data to be labeled. We also show that UNDO robustifies unlearning on the more realistic Weapons of Mass Destruction Proxy (WMDP) benchmark. Since distillation is widely used in practice, incorporating an unlearning step beforehand offers a convenient path to robust capability removal.
Bruce W. Lee, Addie Foote, Alex Infanger, Leni Shor, Harish Kamath, Jacob Goldman-Wetzler, Bryce Woodworth, Alex Cloud, Alexander Matt Turner
NeurIPS1
2025 Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
abstract
As AIs rapidly advance and become more agentic, the risk they pose is governed not only by their capabilities but increasingly by their propensities, including goals and values. Tracking the emergence of goals and values has proven a longstanding problem, and despite much interest over the years it remains unclear whether current AIs have meaningful values. We propose a solution to this problem, leveraging the framework of utility functions to study the internal coherence of AI preferences. Surprisingly, we find that independently-sampled preferences in current LLMs exhibit high degrees of structural coherence, and moreover that this emerges with scale. These findings suggest that value systems emerge in LLMs in a meaningful sense, a finding with broad implications. To study these emergent value systems, we propose utility engineering as a research agenda, comprising both the analysis and control of AI utilities. We uncover problematic and often shocking values in LLM assistants despite existing control measures. These include cases where AIs value themselves over humans and are anti-aligned with specific individuals. To constrain these emergent value systems, we propose methods of utility control. As a case study, we show how aligning utilities with a citizen assembly reduces political biases and generalizes to new scenarios. Whether we like it or not, value systems have already emerged in AIs, and much work remains to fully understand and control these emergent representations.
Mantas Mazeika, Xuwang Yin, Rishub Tamirisa, Jaehyuk Lim, Bruce W. Lee, Richard Ren, Long Phan, Norman Mu, Oliver Zhang, Dan Hendrycks
NeurIPS5
2024 Language Models Don't Learn the Physical Manifestation of Language
abstract
We argue that language-only models don't learn the physical manifestation of language.We present an empirical investigation of visualauditory properties of language through a series of tasks, termed H-TEST.These tasks highlight a fundamental gap between human linguistic understanding and the sensory-deprived linguistic understanding of LLMs.In support of our hypothesis, 1. deliberate reasoning (Chain-of-Thought), 2. few-shot examples, or 3. stronger LLM from the same model family (LLaMA 2 13B → LLaMA 2 70B) has no significant effect on H-TEST performance.We bring in the philosophical case of Mary, who learns about the world in a sensorydeprived environment as a useful conceptual framework to understand how languageonly models learn about the world (Jackson, 1986).Our experiments show that some of the strongest proprietary LLMs stay near random chance baseline accuracy of 50%, highlighting the limitations of linguistic knowledge acquired in the absence of sensory experience.Our code and data are available at .
Bruce W. Lee, Jaehyuk Lim
ACL (1)1
2021 Pushing on Text Readability Assessment: A Transformer Meets Handcrafted Linguistic Features
abstract
We report two essential improvements in readability assessment: 1. three novel features in advanced semantics and 2. the timely evidence that traditional ML models (e.g.Random Forest, using handcrafted features) can combine with transformers (e.g.RoBERTa) to augment model performance.First, we explore suitable transformers and traditional ML models.Then, we extract 255 handcrafted linguistic features using self-developed extraction software.Finally, we assemble those to create several hybrid models, achieving state-of-the-art (SOTA) accuracy on popular datasets in readability assessment.The use of handcrafted features help model performance on smaller datasets.Notably, our RoBERTA-RF-T1 hybrid achieves the near-perfect classification accuracy of 99%, a 20.3% increase from the previous SOTA.
Bruce W. Lee, Yoo Sung Jang, Jason Hyung-Jong Lee
EMNLP (1)1