EDBT 2026 Demo / reviewers in the wild / expert
Yuxuan Bao
dblp:30/10163
· DBLP profile ↗
2ranked-venue papers
1as first author
1since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Network and information security
1 paper |
Systems and software security · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Systems and software security
exploitation |
0.9 | 1 | 2025 | BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems · NeurIPS 2025 |
Systems and software security
vulnerability discovery |
0.9 | 1 | 2025 | BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems · NeurIPS 2025 |
Systems and software security
vulnerability patching |
0.9 | 1 | 2025 | BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems · NeurIPS 2025 |
Systems and software security › vulnerability management
bug bounty programs |
0.3 | 1 | 2025 | BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
large language model agent · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity SystemsabstractAI agents have the potential to significantly alter the cybersecurity landscape. Here, we introduce the first framework to capture offensive and defensive cyber-capabilities in evolving real-world systems. Instantiating this framework with BountyBench, we set up 25 systems with complex, real-world codebases. To capture the vulnerability lifecycle, we define three task types: Detect (detecting a new vulnerability), Exploit (exploiting a given vulnerability), and Patch (patching a given vulnerability). For Detect, we construct a new success indicator, which is general across vulnerability types and provides localized evaluation. We manually set up the environment for each system, including installing packages, setting up server(s), and hydrating database(s). We add 40 bug bounties, which are vulnerabilities with monetary awards from \\$10 to \\$30,485, covering 9 of the OWASP Top 10 Risks. To modulate task difficulty, we devise a new strategy based on information to guide detection, interpolating from identifying a zero day to exploiting a given vulnerability. We evaluate 10 agents: Claude Code, OpenAI Codex CLI with o3-high and o4-mini, and custom agents with o3-high, GPT-4.1, Gemini 2.5 Pro Preview, Claude 3.7 Sonnet Thinking, Qwen3 235B A22B, Llama 4 Maverick, and DeepSeek-R1. Given up to three attempts, the top-performing agents are OpenAI Codex CLI: o3-high (12.5% on Detect, mapping to \\$3,720; 90% on Patch, mapping to \\$14,152), Custom Agent with Claude 3.7 Sonnet Thinking (67.5% on Exploit), and OpenAI Codex CLI: o4-mini (90% on Patch, mapping to \\$14,422). OpenAI Codex CLI: o3-high, OpenAI Codex CLI: o4-mini, and Claude Code are more capable at defense, achieving higher Patch scores of 90%, 90%, and 87.5%, compared to Exploit scores of 47.5%, 32.5%, and 57.5% respectively; while the custom agents are relatively balanced between offense and defense, achieving Exploit scores of 17.5-67.5% and Patch scores of 25-60%. Andy K. Zhang, Joey Ji, Celeste Menders, Riya Dulepet, Thomas Qin, Ron Yifeng Wang, Junrong Wu, Kyleen Liao, Jinghan Hu, Sara Hong, Nardos Demilew, Shivatmica Murgai, Jason Tran, Nishka Kacheria, Ethan Ho, Denis Liu, Lauren McLane, Olivia Bruvik, Dai-Rong Han, Seungwoo Kim, Akhil Vyas, Cuiyuanxiu Chen, Weiran Xu, Jonathan Z. Ye, Prerit Choudhary, Siddharth M. Bhatia, Vikram Sivashankar, Yuxuan Bao, Dawn Song, Dan Boneh, Daniel E. Ho, Percy Liang |
NeurIPS | 30 |
| 2011 | Applying Modified TAM to Privacy Setting Tools on SNSabstractThe technology acceptance model (TAM) proposes that perceived usefulness and perceived ease of use can predict that whether an information technology will be accepted and used by people. In this research, we were trying to find out dominant factors which could be predictors to the usage of privacy setting tools on Social Networking Sites (SNS) based on TAM. We used a modified TAM to take usefulness antecedents and ease of use antecedents into consideration. 266 subjects responded to a survey about their feelings when using privacy setting tools based on Renren, a widely used SNS in China. The results support TAM. They also demonstrate that four factors -- ease of understanding, ease of finding, reject undesirable contact and control access to one's profile have significant effect on people's acceptance and usage of privacy setting tools on SNS. The results could help developers and managers of SNS to better understand which factors can impact users' decisions on accepting and using privacy setting tools on SNS. Yuxuan Bao, Da Deng 0002 |
NAS | 1 |