Chenghua Huang

dblp:352/9764 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 61% Information extraction and text analysis · 23% Language models and text generation · 15%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
policy optimization
0.912025
Token-level Proximal Policy Optimization for Query Generation · EMNLP 2025
Machine learning › Reinforcement learning › policy optimization
proximal policy optimization
0.912025
Token-level Proximal Policy Optimization for Query Generation · EMNLP 2025
Natural language and speech › Language models and text generation › retrieval-augmented generation
query generation
0.912025
Token-level Proximal Policy Optimization for Query Generation · EMNLP 2025
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.912025
Self-Evolved Reward Learning for LLMS · ICLR 2025
Machine learning › Reinforcement learning › reward learning
reward model training
0.912025
Self-Evolved Reward Learning for LLMS · ICLR 2025
Natural language and speech › Information extraction and text analysis
relation extraction
0.712023
Competition or Cooperation? Exploring Unlabeled Data via Challenging Minimax Game for Semi-supervised Relation Extraction · AAAI 2023
Natural language and speech › Information extraction and text analysis › relation extraction
semi-supervised relation extraction
0.712023
Competition or Cooperation? Exploring Unlabeled Data via Challenging Minimax Game for Semi-supervised Relation Extraction · AAAI 2023

Methods — techniques the papers use, named apart from their topics

token-level optimization · 0.9self-training · 0.9self-feedback · 0.9proximal policy optimization · 0.9minimax game · 0.7generator-discriminator · 0.7adversarial training · 0.7
YearPublicationVenuePosition
2025 Token-level Proximal Policy Optimization for Query Generation
abstract
Yichen Ouyang, Lu Wang, Fangkai Yang, Pu Zhao, Chenghua Huang, Jianfeng Liu, Bochen Pang, Yaming Yang, Yuefeng Zhan, Hao Sun, Qingwei Lin, Saravan Rajmohan, Weiwei Deng, Dongmei Zhang, Feng Sun. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yichen Ouyang, Lu Wang 0029, Fangkai Yang, Pu Zhao 0004, Chenghua Huang, Bochen Pang, Yaming Yang 0001, Yuefeng Zhan, Hao Sun 0015, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001, Feng Sun 0008
EMNLP5
2025 Self-Evolved Reward Learning for LLMS
abstract
Reinforcement Learning from Human Feedback (RLHF) is a crucial technique for aligning language models with human preferences and is a key factor in the success of modern conversational models like GPT-4, ChatGPT, and Llama 2. A significant challenge in employing RLHF lies in training a reliable RM, which relies on high-quality labels. Typically, these labels are provided by human experts or a stronger AI, both of which can be costly and introduce bias that may affect the language model's responses. As models improve, human input may become less effective in enhancing their performance. This paper explores the potential of using the RM itself to generate additional training data for a more robust RM. Our experiments demonstrate that reinforcement learning from self-feedback outperforms baseline approaches. We conducted extensive experiments with our approach on multiple datasets, such as HH-RLHF and UltraFeedback, and models including Mistral and Llama 3, comparing it against various baselines. Our results indicate that, even with a limited amount of human-labeled data, learning from self-feedback can robustly enhance the performance of the RM, thereby improving the capabilities of large language models.
Chenghua Huang, Zhizhen Fan, Lu Wang 0029, Fangkai Yang, Pu Zhao 0004, Zeqi Lin, Qingwei Lin, Dongmei Zhang 0001, Saravan Rajmohan, Qi Zhang 0066
ICLR1
2023 Competition or Cooperation? Exploring Unlabeled Data via Challenging Minimax Game for Semi-supervised Relation Extraction
abstract
Semi-Supervised Relation Extraction aims at learning well-performed RE models with limited labeled and large-scale unlabeled data. Existing methods mainly suffer from semantic drift and insufficient supervision, which severely limit the performance. To address these problems, recent work tends to design dual modules to work cooperatively for mutual enhancement. However, the consensus of two modules greatly restricts the model from exploring diverse relation expressions in unlabeled set, which hinders the performance as well as model generalization. To tackle this problem, in this paper, we propose a novel competition-based method AdvSRE. We set up a challenging minimax game on unlabeled data between two modules, Generator and Discriminator, and assign them with conflicting objectives. During the competition game, one module may find any possible chance to beat the other, which develops two modules' abilities until relation expressions cannot be further explored. To exploit label information, Discriminator is further asked to predict specific relation for each sentence. Experiment results on two benchmarks show new state-of-the-art performance over baselines, demonstrating the effectiveness of proposed AdvSRE.
Jianchuan Feng, Chenghua Huang, Zhixu Li, Jianfeng Qu, Yanghua Xiao, Wei Wang 0009
AAAI4