EDBT 2026 Demo / reviewers in the wild / expert
Bufan Gao
dblp:367/5892
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2025
0009-0003-9115-1460ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Trustworthy machine learning · 54% Language models and text generation · 46% | |
| Software engineering, system software, and programming languages
1 paper |
Software testing · 100% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
0.9 | 1 | 2025 | Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases · EMNLP 2025 |
Machine learning › Trustworthy machine learning › fairness › fairness evaluation
gender bias evaluation |
0.9 | 1 | 2025 | Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases · EMNLP 2025 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.9 | 1 | 2025 | Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases · EMNLP 2025 |
Natural language and speech › Language models and text generation › prompting
prompt sensitivity |
0.9 | 1 | 2025 | Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases · EMNLP 2025 |
Software testing › system testing › cyber-physical system testing
autonomous driving system testing |
0.8 | 1 | 2024 | SCTrans: Constructing a Large Public Scenario Dataset for Simulation Testing of Autonomous Driving Systems · ICSE 2024 |
Software testing
simulation-based testing |
0.8 | 1 | 2024 | SCTrans: Constructing a Large Public Scenario Dataset for Simulation Testing of Autonomous Driving Systems · ICSE 2024 |
Software testing
test generation |
0.8 | 1 | 2024 | SCTrans: Constructing a Large Public Scenario Dataset for Simulation Testing of Autonomous Driving Systems · ICSE 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.3 | 1 | 2025 | Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
token-probability metrics · 0.9discrete-choice metrics · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender BiasesabstractAs LLMs are increasingly applied in socially impactful settings, concerns about gender bias have prompted growing efforts both to measure and mitigate such bias.These efforts often rely on evaluation tasks that differ from natural language distributions, as they typically involve carefully constructed task prompts that overtly or covertly signal the presence of gender bias-related content.In this paper, we examine how signaling the evaluative purpose of a task impacts measured gender bias in LLMs.Concretely, we test models under prompt conditions that (1) make the testing context salient, and (2) make gender-focused content salient.We then assess prompt sensitivity across four task formats with both token-probability and discrete-choice metrics.We find that prompts that more clearly align with (gender bias) evaluation framing elicit distinct gender output distributions compared to less evaluation-framed prompts.Discrete-choice metrics further tend to amplify bias relative to probabilistic measures.These findings do not only highlight the brittleness of LLM gender bias evaluations but open a new puzzle for the NLP benchmarking and development community: To what extent can well-controlled testing designs trigger LLM "testing mode" performance, and what does this mean for the ecological validity of future benchmarks. Bufan Gao, Elisa Kreiss |
EMNLP | 1 |
| 2024 | SCTrans: Constructing a Large Public Scenario Dataset for Simulation Testing of Autonomous Driving SystemsabstractFor the safety assessment of autonomous driving systems (ADS), simulation testing has become an important complementary technique to physical road testing. In essence, simulation testing is a scenario-driven approach, whose effectiveness is highly dependent on the quality of given simulation scenarios. Moreover, simulation scenarios should be encoded into well-formatted files, otherwise, ADS simulation platforms cannot take them as inputs. Without large public datasets of simulation scenario files, both industry and academic applications of ADS simulation testing are hindered. Jiarun Dai, Bufan Gao, Mingyuan Luo, Zongan Huang, Zhongrui Li, Yuan Zhang 0009, Min Yang 0002 |
ICSE | 2 |