Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Chenfu Bao

dblp:205/2109 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0000-6484-1552ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Security and privacy · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 45% Reinforcement learning · 28% Trustworthy machine learning · 15%
Software engineering, system software, and programming languages
2 papers
Software maintenance and evolution · 54% Operating systems · 46%
Network and information security
1 paper
Systems and software security · 100%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
1.012026
Safety-Utility Conflicts Are Not Global: Surgical Alignment via Head-Level Diagnosis · ACL (1) 2026
Machine learning › Reinforcement learning
policy optimization
1.012026
Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via sequence-level likelihood · ACL (1) 2026
Machine learning › Trustworthy machine learning › AI safety
safety alignment
1.012026
Safety-Utility Conflicts Are Not Global: Surgical Alignment via Head-Level Diagnosis · ACL (1) 2026
Computer vision › Face, body and person analysis
face forgery detection
0.912025
HAMLET-FFD: Hierarchical Adaptive Multi-modal Learning Embeddings Transformation for Face Forgery Detection · ACM Multimedia 2025
Natural language and speech › Language models and text generation › alignment
preference alignment
0.912025
Indirect Online Preference Optimization via Reinforcement Learning · IJCAI 2025
Natural language and speech › Language models and text generation
preference optimization
0.912025
Indirect Online Preference Optimization via Reinforcement Learning · IJCAI 2025
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.912025
Indirect Online Preference Optimization via Reinforcement Learning · IJCAI 2025
Operating systems › system security › operating system security
kernel security
0.722020
Automatic Hot Patch Generation for Android Kernels · USENIX Security Symposium 2020
Adaptive Android Kernel Live Patching · USENIX Security Symposium 2017
Systems and software security › vulnerability patching
automatic patch generation
0.412020
Automatic Hot Patch Generation for Android Kernels · USENIX Security Symposium 2020
Systems and software security
vulnerability patching
0.412020
Automatic Hot Patch Generation for Android Kernels · USENIX Security Symposium 2020
Software maintenance and evolution › software updates
kernel patching
0.312017
Adaptive Android Kernel Live Patching · USENIX Security Symposium 2017
Software maintenance and evolution › dynamic software updating
live patching
0.312017
Adaptive Android Kernel Live Patching · USENIX Security Symposium 2017
Software maintenance and evolution
software evolution
0.312017
Adaptive Android Kernel Live Patching · USENIX Security Symposium 2017

Methods — techniques the papers use, named apart from their topics

surgical alignment · 1.0sequence-level likelihood · 1.0head-level diagnosis · 1.0group-level rewards · 1.0nash equilibrium · 0.9min-max equilibrium · 0.9hierarchical adaptive multi-modal learning · 0.9embeddings transformation · 0.9adversarial training · 0.9DPO · 0.9hot patching · 0.9live patching · 0.3
YearPublicationVenuePosition
2026 Safety-Utility Conflicts Are Not Global: Surgical Alignment via Head-Level Diagnosis
abstract
Wang Cai, Yilin Wen, Jinchang Hou, Du Su, Guoqiu Wang, Zhonghou Lv, Chenfu Bao, Yunfang Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Wang Cai, Yilin Wen 0007, Jinchang Hou, Du Su, Guoqiu Wang, Zhonghou Lv, Chenfu Bao, Yunfang Wu
ACL (1)7
2026 Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via sequence-level likelihood
abstract
Xingyu Lin, Yilin Wen, Du Su, En Wang, Wenbin Liu, Zhonghou Lv, Jinchang Hou, Chenfu Bao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yilin Wen 0007, Du Su, En Wang, Zhonghou Lv, Jinchang Hou, Chenfu Bao
ACL (1)8
2025 Indirect Online Preference Optimization via Reinforcement Learning
abstract
Human preference alignment (HPA) aims to ensure Large Language Models (LLMs) responding appropriately to meet human moral and ethical requirements. Existing methods, such as RLHF and DPO, rely heavily on high-quality human annotation, which restrict the efficiency of iterative online model refinement. To address the inefficiencies of human annotation acquisition, iterated online strategy advocates the use of fine-tuned LLMs to self-generate preference data. However, this approach is prone to distribution bias, because of differences between human and model annotations, as well as modeling errors between simulators and real-world contexts. To mitigate the impact of distribution bias, we adopt the principles of adversarial training, framing a zero-sum two-player game with a protagonist agent and an adversarial agent. With the adversarial agent challenging the alignment of protagonist agent, we continuously refine the protagonist’s performance. By utilizing min-max equilibrium and Nash equilibrium strategies, we propose Indirect Online Preference Optimization (IOPO) mechanism that enables the protagonist agent to converge without bias while maintaining linear computational complexity. Extensive experiments across three real-world datasets demonstrate that IOPO outperforms state-of-the-art alignment methods in both offline and online scenarios, evidenced by standard alignment metrics and human evaluations. This innovation reduces the time required for model iterations from months to one week, alleviates distribution shifts, and significantly cuts annotation costs.
En Wang, Du Su, Chenfu Bao, Zhonghou Lv, Funing Yang, Yuanbo Xu
IJCAI4
2025 HAMLET-FFD: Hierarchical Adaptive Multi-modal Learning Embeddings Transformation for Face Forgery Detection
Jialei Cui, Jianwei Du, Chenfu Bao
ACM Multimedia6
2020 Automatic Hot Patch Generation for Android Kernels
Zhengzi Xu, Longri Zheng, Liangzhao Xia, Chenfu Bao, Zhi Wang 0004, Yang Liu 0003
USENIX Security Symposium5
2017 Adaptive Android Kernel Live Patching
Zhi Wang 0004, Liangzhao Xia, Chenfu Bao, Tao Wei 0002
USENIX Security Symposium5