EDBT 2026 Demo / reviewers in the wild / expert
Xiaoyan Bai
dblp:63/3140
· DBLP profile ↗
5ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Efficient and distributed learning · 28% Information extraction and text analysis · 16% Trustworthy machine learning · 15% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning |
0.9 | 1 | 2025 | Concept Incongruence: An Exploration of Time and Death in Role Playing · NeurIPS 2025 |
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | Concept Incongruence: An Exploration of Time and Death in Role Playing · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › model compression › sparsity
activation sparsity |
0.8 | 1 | 2024 | Learn To be Efficient: Build Structured Sparsity in Large Language Models · NeurIPS 2024 |
Natural language and speech › Machine translation › statistical machine translation
alignment models |
0.8 | 1 | 2024 | A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity · ICML 2024 |
Machine learning › Efficient and distributed learning › inference efficiency
LLM inference optimization |
0.8 | 1 | 2024 | Learn To be Efficient: Build Structured Sparsity in Large Language Models · NeurIPS 2024 |
Machine learning › Efficient and distributed learning
model compression |
0.8 | 1 | 2024 | Learn To be Efficient: Build Structured Sparsity in Large Language Models · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
toxicity reduction |
0.8 | 1 | 2024 | A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity · ICML 2024 |
Natural language and speech › Information extraction and text analysis
keyphrase extraction |
0.7 | 1 | 2023 | PromptRank: Unsupervised Keyphrase Extraction Using Prompt · ACL (1) 2023 |
Computer vision › Vision and language › vision-language model
prompt learning |
0.7 | 1 | 2023 | PromptRank: Unsupervised Keyphrase Extraction Using Prompt · ACL (1) 2023 |
Natural language and speech › Information extraction and text analysis › keyphrase extraction
unsupervised keyphrase extraction |
0.7 | 1 | 2023 | PromptRank: Unsupervised Keyphrase Extraction Using Prompt · ACL (1) 2023 |
Machine learning › Trustworthy machine learning
interpretability |
0.5 | 2 | 2025 | Concept Incongruence: An Exploration of Time and Death in Role Playing · NeurIPS 2025 A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity · ICML 2024 |
Machine learning › Representation and self-supervised learning
probing |
0.3 | 1 | 2025 | Concept Incongruence: An Exploration of Time and Death in Role Playing · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
probing · 0.9behavioral metrics · 0.9training algorithm for sparsity · 0.8mechanistic analysis · 0.8hardware-aware kernel · 0.8direct preference optimization (DPO) · 0.8prompt-based ranking · 0.7encoder-decoder language model · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Concept Incongruence: An Exploration of Time and Death in Role PlayingabstractConsider this prompt "Draw a unicorn with two horns". Should large language models (LLMs) recognize that a unicorn has only one horn by definition and ask users for clarifications, or proceed to generate something anyway? We introduce *concept incongruence* to capture such phenomena where concept boundaries clash with each other, either in user prompts or in model representations, often leading to under-specified or mis-specified behaviors. In this work, we take the first step towards defining and analyzing model behavior under concept incongruence. Focusing on temporal boundaries in the Role-Play setting, we propose three behavioral metrics---abstention rate, conditional accuracy, and answer rate---to quantify model behavior under incongruence due to the role's death. We show that models fail to abstain after death and suffer from an accuracy drop compared to the Non-Role-Play setting. Through probing experiments, we identify two main causes: (i) unreliable encoding of the "death" state across different years, leading to unsatisfactory abstention behavior, and (ii) role playing causes shifts in the model’s temporal representations, resulting in accuracy drops. We leverage these insights to improve consistency in the model's abstention and answer behaviors. Our findings suggest that concept incongruence leads to unexpected model behaviors and point to future directions on improving model behavior under concept incongruence. Xiaoyan Bai, Ike Peng, Chenhao Tan |
NeurIPS | 1 |
| 2024 | A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and ToxicityabstractWhile alignment algorithms are commonly used to tune pre-trained language models towards user preferences, we lack explanations for the underlying mechanisms in which models become ``aligned'', thus making it difficult to explain phenomena like jailbreaks. In this work we study a popular algorithm, direct preference optimization (DPO), and the mechanisms by which it reduces toxicity. Namely, we first study how toxicity is represented and elicited in pre-trained language models (GPT2-medium, Llama2-7b). We then apply DPO with a carefully crafted pairwise dataset to reduce toxicity. We examine how the resulting models avert toxic outputs, and find that capabilities learned from pre-training are not removed, but rather bypassed. We use this insight to demonstrate a simple method to un-align the models, reverting them back to their toxic behavior. Andrew Lee 0001, Xiaoyan Bai, Itamar Pres, Martin Wattenberg, Jonathan K. Kummerfeld, Rada Mihalcea |
ICML | 2 |
| 2024 | Learn To be Efficient: Build Structured Sparsity in Large Language ModelsabstractLarge Language Models (LLMs) have achieved remarkable success with their billion-level parameters, yet they incur high inference overheads. The emergence of activation sparsity in LLMs provides a natural approach to reduce this cost by involving only parts of the parameters for inference. However, existing methods only focus on utilizing this naturally formed activation sparsity in a post-training setting, overlooking the potential for further amplifying this inherent sparsity. In this paper, we hypothesize that LLMs can learn to be efficient by achieving more structured activation sparsity. To achieve this, we introduce a novel training algorithm, Learn-To-be-Efficient (LTE), designed to train efficiency-aware LLMs to learn to activate fewer neurons and achieve a better trade-off between sparsity and performance. Furthermore, unlike SOTA MoEfication methods, which mainly focus on ReLU-based models, LTE can also be applied to LLMs like LLaMA using non-ReLU activations. Extensive evaluation on language understanding, language generation, and instruction tuning tasks show that LTE consistently outperforms SOTA baselines. Along with our hardware-aware custom kernel implementation, LTE reduces LLaMA2-7B inference latency by 25% at 50% sparsity. Haizhong Zheng, Xiaoyan Bai, Xueshen Liu, Z. Morley Mao, Beidi Chen, Fan Lai 0001, Atul Prakash 0001 |
NeurIPS | 2 |
| 2023 | PromptRank: Unsupervised Keyphrase Extraction Using PromptabstractThe keyphrase extraction task refers to the automatic selection of phrases from a given document to summarize its core content.Stateof-the-art (SOTA) performance has recently been achieved by embedding-based algorithms, which rank candidates according to how similar their embeddings are to document embeddings.However, such solutions either struggle with the document and candidate length discrepancies or fail to fully utilize the pretrained language model (PLM) without further fine-tuning.To this end, in this paper, we propose a simple yet effective unsupervised approach, PromptRank, based on the PLM with an encoder-decoder architecture.Specifically, PromptRank feeds the document into the encoder and calculates the probability of generating the candidate with a designed prompt by the decoder.We extensively evaluate the proposed PromptRank on six widely used benchmarks.PromptRank outperforms the SOTA approach MDERank, improving the F 1 score relatively by 34.18%, 24.87%, and 17.57% for 5, 10, and 15 returned results, respectively.This demonstrates the great potential of using prompt for unsupervised keyphrase extraction.We release our code at this url. Aobo Kong, Shiwan Zhao, Qicheng Li, Xiaoyan Bai |
ACL (1) | 7 |
| 2011 | Cross-Cultural Learning Design: Past the Head and the Hands to the HEART of the MatterabstractDeveloping countries are looking to satisfy ever-growing demand for education by leveraging the access and communication capabilities afforded by e-learning and the increasing range of open education resources made available through universities and development organisations globally. However a 'one size fits all', predominantly Western approach to developing and implementing educational initiatives has often failed in information and communication technology for development (ICT4D) projects. Over the past few years, we have been piloting and refining a learning design support strategy, entitled HEART: HEaring And Realizing Teaching-voice. HEART aims to enhance educators' learning design awareness and capability by eliciting and depicting the pedagogical beliefs underpinning a course or learning design. HEART has shown its potential to promote discussion and reflection of the beliefs and values that underpin course design. These beliefs can otherwise act as invisible but powerful obstacles to achieving the cross-cultural understandings necessary for effective e-learning design and implementation. This workshop will enable each participant to use the HEART visualization strategy to elicit and depict cultural-pedagogical beliefs relating to a teaching project or learning context of personal interest, demonstrating its ability to facilitate more effective cross-cultural learning design practice. Adam Blake, Claire Donald, Ashwini Datt, Xiaoyan Bai |
ICALT | 4 |