Ke Yang 0003

dblp:80/4136-3 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0002-1617-5986ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Trustworthy machine learning · 69% Language models and text generation · 22% Reinforcement learning · 9%
Theoretical computer science
1 paper
Algorithmic game theory and mechanism design · 77% Mathematical optimization · 23%

Topics — the 14 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
2.642024
Bias and Volatility: A Statistical Framework for Evaluating Large Language Model's Stereotypes and the Associated Generation Inconsistency · NeurIPS 2024
Non-Invasive Fairness in Learning Through the Lens of Data Drift · ICDE 2024
ADEPT: A DEbiasing PrompT Framework · AAAI 2023
Natural language and speech › Language models and text generation
LLM agents
0.912025
AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents · ICLR 2025
Natural language and speech › Language models and text generation › LLM agents
web agents
0.912025
AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents · ICLR 2025
Machine learning › Trustworthy machine learning › fairness › fairness evaluation
stereotype bias evaluation
0.812024
Bias and Volatility: A Statistical Framework for Evaluating Large Language Model's Stereotypes and the Associated Generation Inconsistency · NeurIPS 2024
Machine learning › Trustworthy machine learning › fairness › bias mitigation
large language model debiasing
0.712023
ADEPT: A DEbiasing PrompT Framework · AAAI 2023
Machine learning › Trustworthy machine learning › fairness
within-group fairness
0.412019
Balanced Ranking with Diversity Constraints · IJCAI 2019
Algorithmic game theory and mechanism design › social choice › computational social choice
fair ranking
0.412019
Balanced Ranking with Diversity Constraints · IJCAI 2019
Machine learning › Trustworthy machine learning › fairness
ranking fairness
0.312018
A Nutritional Label for Rankings · SIGMOD Conference 2018
Machine learning › Trustworthy machine learning › fairness › bias in language models
dialect bias
0.312025
A Multi-Agent Framework for Mitigating Dialect Biases in Privacy Policy Question-Answering Systems · ACL (1) 2025
Natural language and speech › Language models and text generation
alignment
0.212024
Bias and Volatility: A Statistical Framework for Evaluating Large Language Model's Stereotypes and the Associated Generation Inconsistency · NeurIPS 2024
Natural language and speech › Language models and text generation
prompt tuning
0.212023
ADEPT: A DEbiasing PrompT Framework · AAAI 2023
Mathematical optimization
integer programming
0.112019
Balanced Ranking with Diversity Constraints · IJCAI 2019
Information retrieval
ranking
0.112018
A Nutritional Label for Rankings · SIGMOD Conference 2018
Information retrieval › ranking › learning to rank › robust ranking
ranking stability
0.112018
A Nutritional Label for Rankings · SIGMOD Conference 2018

Methods — techniques the papers use, named apart from their topics

zero-shot prompting · 0.9observation and action space alignment · 0.9multi-agent framework · 0.9statistical framework · 0.8reweighing · 0.8model splitting · 0.8diffair · 0.8conformance constraints · 0.8confair · 0.8bias-variance decomposition · 0.8integer linear programming · 0.4nutritional label · 0.3
YearPublicationVenuePosition
2026 TRACER: Early Failure Detection for Task-Oriented Dialogue
abstract
Task-oriented dialogue systems often fail before the final breakdown is obvious, but most evaluation only measures failure after the conversation has already gone wrong. We present TRACER, a method for early failure detection in task-oriented dialogue. TRACER predicts from a partial dialogue whether the full conversation will eventually fail by combining simple trajectory signals from belief-state changes with text representations of the evolving dialogue state. We evaluate the method in both oracle and generated belief-state settings, and test how well it works when only 25%, 50%, 75%, or 100% of the dialogue is visible. Across these settings, TRACER detects useful failure signals well before the end of the conversation and outperforms heuristic, classical, and single-stream baselines. These results suggest that early failure detection can provide a practical warning signal for dialogue systems before the interaction fully breaks down. Source code can be found here: https://github.com/erfan-nourbakhsh/TRACER.
Erfan Nourbakhsh, Rocky Slavin, Ke Yang 0003, Anthony Rios
SIGDIAL3
2025 A Multi-Agent Framework for Mitigating Dialect Biases in Privacy Policy Question-Answering Systems
abstract
Đorđe Klisura, Astrid R Bernaga Torres, Anna Karen Gárate-Escamilla, Rajesh Roshan Biswal, Ke Yang, Hilal Pataci, Anthony Rios. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Dorde Klisura, Astrid R. Bernaga Torres, Anna Karen Gárate-Escamilla, Rajesh Roshan Biswal, Ke Yang 0003, Hilal Pataci, Anthony Rios
ACL (1)5
2025 Persona-DB: Efficient Large Language Model Personalization for Response Prediction with Collaborative Data Refinement
abstract
The increasing demand for personalized interactions with large language models (LLMs) calls for methodologies capable of accurately and efficiently identifying user opinions and preferences. Retrieval augmentation emerges as an effective strategy, as it can accommodate a vast number of users without the costs from fine-tuning. Existing research, however, has largely focused on enhancing the retrieval stage and devoted limited exploration toward optimizing the representation of the database, a crucial aspect for tasks such as personalization. In this work, we examine the problem from a novel angle, focusing on how data can be better represented for more data-efficient retrieval in the context of LLM customization. To tackle this challenge, we introduce Persona-DB, a simple yet effective framework consisting of a hierarchical construction process to improve generalization across task contexts and collaborative refinement to effectively bridge knowledge gaps among users. In the evaluation of response prediction, Persona-DB demonstrates superior context efficiency in maintaining accuracy with a significantly reduced retrieval size, a critical advantage in scenarios with extensive histories or limited context windows. Our experiments also indicate a marked improvement of over 10% under cold-start scenarios, when users have extremely sparse data. Furthermore, our analysis reveals the increasing importance of collaborative knowledge as the retrieval capacity expands.
Chenkai Sun, Ke Yang 0003, Revanth Gangi Reddy, Yi R. Fung 0001, Hou Pong Chan, Kevin Small, ChengXiang Zhai, Heng Ji 0001
COLING2
2025 AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents
abstract
Autonomy via agents based on large language models (LLMs) that can carry out personalized yet standardized tasks presents a significant opportunity to drive human efficiency. There is an emerging need and interest in automating web tasks (e.g., booking a hotel for a given date within a budget). Being a practical use case itself, the web agent also serves as an important proof-of-concept example for various agent grounding scenarios, with its success promising advancements in many future applications. Meanwhile, much prior research focuses on handcrafting their web agent strategies (e.g., agent's prompting templates, reflective workflow, role-play and multi-agent systems, search or sampling methods, etc.) and the corresponding in-context examples. However, these custom strategies often struggle with generalizability across all potential real-world applications. On the other hand, there has been limited study on the misalignment between a web agent's observation and action representation, and the data on which the agent's underlying LLM has been pre-trained. This discrepancy is especially notable when LLMs are primarily trained for language completion rather than tasks involving embodied navigation actions and symbolic web elements. In our study, we enhance an LLM-based web agent by simply refining its observation and action space, aligning these more closely with the LLM's capabilities. This approach enables our base agent to significantly outperform previous methods on a wide variety of web tasks. Specifically, on WebArena, a benchmark featuring general-purpose web interaction tasks, our agent AgentOccam surpasses the previous state-of-the-art and concurrent work by 9.8 (+29.4%) and 5.9 (+15.8%) absolute points respectively, and boosts the success rate by 26.6 points (+161%) over similar plain web agents with its observation and action space alignment. Furthermore, on WebVoyager benchmark comprising tasks defined on real-world websites, AgentOccam exceeds the former best agent by 2.4 points (+4.6%) on tasks with deterministic answers. We achieve this without using in-context examples, new agent roles, online feedback or search strategies. AgentOccam's simple design highlights LLMs' impressive zero-shot performance on web tasks, and underlines the critical role of carefully tuning observation and action spaces for LLM-based agents.
Ke Yang 0003, Yao Liu 0009, Sapana Chaudhary, Rasool Fakoor, Pratik Chaudhari, George Karypis, Huzefa Rangwala
ICLR1
2024 Non-Invasive Fairness in Learning Through the Lens of Data Drift
abstract
Machine Learning models are widely employed to drive many modern data systems. While they are undeniably powerful tools, ML models often demonstrate imbalanced performance and unfair behaviors. The root of this problem often lies in the fact that different subpopulations commonly display divergent trends: as a learning algorithm tries to identify trends in the data, it naturally favors the trends of the majority groups, leading to a model that performs poorly and unfairly for minority populations. Our goal is to improve the fairness and trustworthiness of ML models by applying only non-invasive interventions, which don't alter the data or the learning algorithm. We use a simple but key insight: the divergence of trends between different popu-lations, and, consecutively, between a learned model and minority populations, is analogous to data drift, which indicates poor conformance between parts of the data and the trained model. We explore two strategies (model-splitting and reweighing) to resolve this drift, aiming to improve the overall conformance of models to the underlying data. Both our methods introduce novel ways to employ the recently-proposed data profiling primitive of Conformance Constraints. Our splitting approach is based on a simple data drift strategy: training separate models for different populations. Our DifFair algorithm enhances this simple strategy by employing conformance constraints, learned over the data partitions, to select the appropriate model to use for predictions on each serving tuple. However, the performance of such a multi-model strategy can degrade severely under poor representation of some groups in the data. We thus propose a single-model, reweighing strategy, ConFair, to overcome this limitation. ConFair employs conformance constraints in a novel way to derive weights for training data, which are then used to build a single model. Our experimental evaluation over 7 real-world datasets shows that both DifFair and ConFair improve the fairness of ML models. We demonstrate scenarios where DifFair has an edge, though ConFair has the greatest practical impact and outperforms other baselines. Moreover, as a model-agnostic technique, ConFairstays robust when used against different models than the ones on which the weights have been learned, which is not the case for other states of the art.
Ke Yang 0003, Alexandra Meliou
ICDE1
2024 Bias and Volatility: A Statistical Framework for Evaluating Large Language Model's Stereotypes and the Associated Generation Inconsistency
abstract
We present a novel statistical framework for analyzing stereotypes in large language models (LLMs) by systematically estimating the bias and variation in their generation. Current evaluation metrics in the alignment literature often overlook the randomness of stereotypes caused by the inconsistent generative behavior of LLMs. For example, this inconsistency can result in LLMs displaying contradictory stereotypes, including those related to gender or race, for identical professions across varied contexts. Neglecting such inconsistency could lead to misleading conclusions in alignment evaluations and hinder the accurate assessment of the risk of LLM applications perpetuating or amplifying social stereotypes and unfairness.This work proposes a Bias-Volatility Framework (BVF) that estimates the probability distribution function of LLM stereotypes. Specifically, since the stereotype distribution fully captures an LLM's generation variation, BVF enables the assessment of both the likelihood and extent to which its outputs are against vulnerable groups, thereby allowing for the quantification of the LLM's aggregated discrimination risk. Furthermore, we introduce a mathematical framework to decompose an LLM’s aggregated discrimination risk into two components: bias risk and volatility risk, originating from the mean and variation of LLM’s stereotype distribution, respectively. We apply BVF to assess 12 commonly adopted LLMs and compare their risk levels. Our findings reveal that: i) Bias risk is the primary cause of discrimination risk in LLMs; ii) Most LLMs exhibit significant pro-male stereotypes for nearly all careers; iii) Alignment with reinforcement learning from human feedback lowers discrimination by reducing bias, but increases volatility; iv) Discrimination risk in LLMs correlates with key sociol-economic factors like professional salaries. Finally, we emphasize that BVF can also be used to assess other dimensions of generation inconsistency's impact on LLM behavior beyond stereotypes, such as knowledge mastery.
Ke Yang 0003, Zehan Qi, Yang Yu 0011, ChengXiang Zhai
NeurIPS2
2024 DEEILS: Data Ethics Embedded Interactive Learning System for Computer Science Students
abstract
In the rapidly evolving computing technology landscape, ethical issues arise at every stage of the data lifecycle, from collection to downstream predictive analytics, including but not limited to pre-existing bias, privacy, fairness, and accountability of machine learning and artificial intelligence algorithms. Data ethics education for students in Computer Science and related STEM programs has become a focal point of discussion and innovation. This work introduces a system, DEEILS, that allows students to learn about ethical issues at each stage of the data lifecycle through multi-media interactive modules and real-world scenario simulations. It also supports instructors with customized components for different levels of courses. DEEILS consists of three phases: .Collection, Preparation, and Analytics, which represent the stages of how data is created and evolves through its lifecycle. Within each module, DEEILS integrates components to simulate the steps that data practitioners take in real-world scenarios, such as cleaning, de-identification, and feature engineering in the Preparation phase. It then guides students through the ethical issues that can arise in each component, using examples of real datasets and commonly used computational techniques for that component. Through interactive content with real-world applications, DEEILS aims to provide an adaptive, immersive learning environment to facilitate education on ethical issues in data-driven science.
Ke Yang 0003
SIGCSE (2)1
2023 ADEPT: A DEbiasing PrompT Framework
abstract
Several works have proven that finetuning is an applicable approach for debiasing contextualized word embeddings. Similarly, discrete prompts with semantic meanings have shown to be effective in debiasing tasks. With unfixed mathematical representation at the token level, continuous prompts usually surpass discrete ones at providing a pre-trained language model (PLM) with additional task-specific information. Despite this, relatively few efforts have been made to debias PLMs by prompt tuning with continuous prompts compared to its discrete counterpart. Furthermore, for most debiasing methods that alter a PLM's original parameters, a major problem is the need to not only decrease the bias in the PLM but also to ensure that the PLM does not lose its representation ability. Finetuning methods typically have a hard time maintaining this balance, as they tend to violently remove meanings of attribute words (like the words developing our concepts of "male" and "female" for gender), which also leads to an unstable and unpredictable training process. In this paper, we propose ADEPT, a method to debias PLMs using prompt tuning while maintaining the delicate balance between removing biases and ensuring representation ability. To achieve this, we propose a new training criterion inspired by manifold learning and equip it with an explicit debiasing term to optimize prompt tuning. In addition, we conduct several experiments with regard to the reliability, quality, and quantity of a previously proposed attribute training corpus in order to obtain a clearer prototype of a certain attribute, which indicates the attribute's position and relative distances to other words on the manifold. We evaluate ADEPT on several widely acknowledged debiasing benchmarks and downstream tasks, and find that it achieves competitive results while maintaining (and in some cases even improving) the PLM's representation ability. We further visualize words' correlation before and after debiasing a PLM, and give some possible explanations for the visible effects.
Ke Yang 0003, Charles Yu, Yi R. Fung 0001, Manling Li, Heng Ji 0001
AAAI1
2019 Balanced Ranking with Diversity Constraints
abstract
Many set selection and ranking algorithms have recently been enhanced with diversity constraints that aim to explicitly increase representation of historically disadvantaged populations, or to improve the over-all representativeness of the selected set. An unintended consequence of these constraints, however, is reduced in-group fairness: the selected candidates from a given group may not be the best ones, and this unfairness may not be well-balanced across groups. In this paper we study this phenomenon using datasets that comprise multiple sensitive attributes. We then introduce additional constraints, aimed at balancing the in-group fairness across groups, and formalize the induced optimization problems as integer linear programs. Using these programs, we conduct an experimental evaluation with real datasets, and quantify the feasible trade-offs between balance and overall performance in the presence of diversity constraints.
Ke Yang 0003, Vasilis Gkatzelis, Julia Stoyanovich
IJCAI1
2018 Online Set Selection with Fairness and Diversity Constraints
Julia Stoyanovich, Ke Yang 0003, H. V. Jagadish
EDBT2
2018 A Nutritional Label for Rankings
abstract
Algorithmic decisions often result in scoring and ranking individuals to determine credit worthiness, qualifications for college admissions and employment, and compatibility as dating partners. While automatic and seemingly objective, ranking algorithms can discriminate against individuals and protected groups, and exhibit low diversity. Furthermore, ranked results are often unstable -- small changes in the input data or in the ranking methodology may lead to drastic changes in the output, making the result uninformative and easy to manipulate. Similar concerns apply in cases where items other than individuals are ranked, including colleges, academic departments, or products. Despite the ubiquity of rankers, there is, to the best of our knowledge, no technical work that focuses on making rankers transparent.
Ke Yang 0003, Julia Stoyanovich, Abolfazl Asudeh, Bill Howe, H. V. Jagadish, Gerome Miklau
SIGMOD Conference1
2017 Measuring Fairness in Ranked Outputs
abstract
Ranking and scoring are ubiquitous. We consider the setting in which an institution, called a ranker, evaluates a set of individuals based on demographic, behavioral or other characteristics. The final output is a ranking that represents the relative quality of the individuals. While automatic and therefore seemingly objective, rankers can, and often do, discriminate against individuals and systematically disadvantage members of protected groups. This warrants a careful study of the fairness of a ranking scheme, to enable data science for social good applications, among others.
Ke Yang 0003, Julia Stoyanovich
SSDBM1