Bufan Gao

dblp:367/5892 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
0009-0003-9115-1460ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 54% Language models and text generation · 46%
Software engineering, system software, and programming languages
1 paper
Software testing · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
0.912025
Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases · EMNLP 2025
Machine learning › Trustworthy machine learning › fairness › fairness evaluation
gender bias evaluation
0.912025
Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases · EMNLP 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases · EMNLP 2025
Natural language and speech › Language models and text generation › prompting
prompt sensitivity
0.912025
Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases · EMNLP 2025
Software testing › system testing › cyber-physical system testing
autonomous driving system testing
0.812024
SCTrans: Constructing a Large Public Scenario Dataset for Simulation Testing of Autonomous Driving Systems · ICSE 2024
Software testing
simulation-based testing
0.812024
SCTrans: Constructing a Large Public Scenario Dataset for Simulation Testing of Autonomous Driving Systems · ICSE 2024
Software testing
test generation
0.812024
SCTrans: Constructing a Large Public Scenario Dataset for Simulation Testing of Autonomous Driving Systems · ICSE 2024
Machine learning › Trustworthy machine learning
interpretability
0.312025
Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

token-probability metrics · 0.9discrete-choice metrics · 0.9
YearPublicationVenuePosition
2025 Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases
abstract
As LLMs are increasingly applied in socially impactful settings, concerns about gender bias have prompted growing efforts both to measure and mitigate such bias.These efforts often rely on evaluation tasks that differ from natural language distributions, as they typically involve carefully constructed task prompts that overtly or covertly signal the presence of gender bias-related content.In this paper, we examine how signaling the evaluative purpose of a task impacts measured gender bias in LLMs.Concretely, we test models under prompt conditions that (1) make the testing context salient, and (2) make gender-focused content salient.We then assess prompt sensitivity across four task formats with both token-probability and discrete-choice metrics.We find that prompts that more clearly align with (gender bias) evaluation framing elicit distinct gender output distributions compared to less evaluation-framed prompts.Discrete-choice metrics further tend to amplify bias relative to probabilistic measures.These findings do not only highlight the brittleness of LLM gender bias evaluations but open a new puzzle for the NLP benchmarking and development community: To what extent can well-controlled testing designs trigger LLM "testing mode" performance, and what does this mean for the ecological validity of future benchmarks.
Bufan Gao, Elisa Kreiss
EMNLP1
2024 SCTrans: Constructing a Large Public Scenario Dataset for Simulation Testing of Autonomous Driving Systems
abstract
For the safety assessment of autonomous driving systems (ADS), simulation testing has become an important complementary technique to physical road testing. In essence, simulation testing is a scenario-driven approach, whose effectiveness is highly dependent on the quality of given simulation scenarios. Moreover, simulation scenarios should be encoded into well-formatted files, otherwise, ADS simulation platforms cannot take them as inputs. Without large public datasets of simulation scenario files, both industry and academic applications of ADS simulation testing are hindered.
Jiarun Dai, Bufan Gao, Mingyuan Luo, Zongan Huang, Zhongrui Li, Yuan Zhang 0009, Min Yang 0002
ICSE2