Honglin Guo

dblp:60/10205 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Language models and text generation · 44% Reinforcement learning · 21% Generative modeling · 12%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 100%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 17 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
LLM agents
1.922026
AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments · ACL (1) 2026
AgentGym: Evaluating and Training Large Language Model-based Agents across Diverse Environments · ACL (1) 2025
Natural language and speech › Language models and text generation
agent benchmarking
1.012026
AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments · ACL (1) 2026
Natural language and speech › Language models and text generation
code language models
1.012026
OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding · ACL (1) 2026
Natural language and speech › Language models and text generation
alignment
0.912025
Pre-Trained Policy Discriminators are General Reward Models · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.912025
TRCE: Towards Reliable Malicious Concept Erasure in Text-to-Image Diffusion Models · ICCV 2025
Machine learning › Trustworthy machine learning
interpretability
0.912025
BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset · NeurIPS 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset · NeurIPS 2025
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.912025
Pre-Trained Policy Discriminators are General Reward Models · NeurIPS 2025
Machine learning › Reinforcement learning › reward learning
reward modeling
0.912025
Pre-Trained Policy Discriminators are General Reward Models · NeurIPS 2025
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model
0.912025
TRCE: Towards Reliable Malicious Concept Erasure in Text-to-Image Diffusion Models · ICCV 2025
Data integration and cleaning
data quality
0.912025
CritiQ: Mining Data Quality Criteria from Human Preferences · ACL (1) 2025
Machine learning › Reinforcement learning
curriculum reinforcement learning
0.812024
Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning · ICML 2024
Natural language and speech › Language models and text generation
large language model reasoning
0.812024
Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning · ICML 2024
Machine learning › Learning paradigms › curriculum learning
reverse curriculum learning
0.812024
Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning · ICML 2024
Machine learning › Reinforcement learning
policy optimization
0.312025
Pre-Trained Policy Discriminators are General Reward Models · NeurIPS 2025
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement fine-tuning
0.312025
Pre-Trained Policy Discriminators are General Reward Models · NeurIPS 2025
Computational science and engineering
benchmark dataset
0.312025
BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

process-based verification · 1.7preference mining · 1.7malicious concept erasure · 1.7human-in-the-loop curation · 1.7benchmarking · 1.0scaling law analysis · 0.9pre-training · 0.9policy discriminator · 0.9policy gradient · 0.8outcome supervision · 0.8
YearPublicationVenuePosition
2026 OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding
abstract
Deming Ding, Shichun Liu, Enhui Yang, Jiahang Lin, Ziying Chen, Shihan Dou, Honglin Guo, Weiyu Cheng, Pengyu Zhao, Chengjun Xiao, Qunhong Zeng, Qi Zhang, Xuanjing Huang, Qidi Xu, Tao Gui. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Deming Ding, Shichun Liu, Enhui Yang, Jiahang Lin, Ziying Chen, Shihan Dou, Honglin Guo, Weiyu Cheng, Chengjun Xiao, Qunhong Zeng, Qi Zhang 0001, Xuanjing Huang 0001, Qidi Xu, Tao Gui
ACL (1)7
2026 AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments
abstract
Zhiheng Xi, Dingwen Yang, Jiaqi Liu, Jixuan Huang, Honglin Guo, Baodai Huang, Tinggang Chen, Qi Zhang, Zhonghang Lu, Chenyu Liu, Jiajun Sun, Jiazheng Zhang, Dingwei Zhu, Xin Guo, Junzhe Wang, Zhihao Zhang, Yuming Yang, Junjie Ye, Minghe Gao, Dongrui Liu, Jiaming Ji, Guohao Li, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhiheng Xi, Dingwen Yang, Jixuan Huang, Honglin Guo, Baodai Huang, Tinggang Chen, Qi Zhang 0001, Zhonghang Lu, Jiazheng Zhang, Dingwei Zhu, Junzhe Wang 0001, Zhihao Zhang 0002, Yuming Yang 0001, Junjie Ye 0005, Minghe Gao, Dongrui Liu, Jiaming Ji, Tao Gui, Xuanjing Huang 0001
ACL (1)5
2025 CritiQ: Mining Data Quality Criteria from Human Preferences
abstract
Honglin Guo, Kai Lv, Qipeng Guo, Tianyi Liang, Zhiheng Xi, Demin Song, Qiuyinzhe Zhang, Yu Sun, Kai Chen, Xipeng Qiu, Tao Gui. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Honglin Guo, Kai Lv 0001, Qipeng Guo, Tianyi Liang 0002, Zhiheng Xi, Demin Song, Qiuyinzhe Zhang, Yu Sun 0031, Kai Chen 0026, Xipeng Qiu, Tao Gui
ACL (1)1
2025 AgentGym: Evaluating and Training Large Language Model-based Agents across Diverse Environments
abstract
Zhiheng Xi, Yiwen Ding, Wenxiang Chen, Boyang Hong, Honglin Guo, Junzhe Wang, Xin Guo, Dingwen Yang, Chenyang Liao, Wei He, Songyang Gao, Lu Chen, Rui Zheng, Yicheng Zou, Tao Gui, Qi Zhang, Xipeng Qiu, Xuanjing Huang, Zuxuan Wu, Yu-Gang Jiang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhiheng Xi, Yiwen Ding, Wenxiang Chen, Boyang Hong, Honglin Guo, Junzhe Wang 0001, Dingwen Yang, Chenyang Liao, Wei He 0024, Songyang Gao, Lu Chen 0001, Yicheng Zou, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001, Zuxuan Wu, Yu-Gang Jiang 0001
ACL (1)5
2025 TRCE: Towards Reliable Malicious Concept Erasure in Text-to-Image Diffusion Models
Rui-dong Chen, Honglin Guo, Lanjun Wang, Chenyu Zhang 0003, Weizhi Nie, Anan Liu
ICCV2
2025 Pre-Trained Policy Discriminators are General Reward Models
abstract
We offer a novel perspective on reward modeling by formulating it as a policy discriminator, which quantifies the difference between two policies to generate a reward signal, guiding the training policy towards a target policy with desired behaviors. Based on this conceptual insight, we propose a scalable pre-training method named POLicy DiscriminAtive LeaRning (POLAR), which trains a reward model (RM) to discern identical policies and discriminate different ones. Unlike traditional reward modeling methods relying on absolute preferences, POLAR captures the relative difference between one policy and an arbitrary target policy, which is a scalable, high-level optimization objective suitable for modeling generic ranking relationships. Leveraging the POLAR pre-training paradigm, we present a series of RMs with parameter scales from 1.8B to 7B. Empirical results show that POLAR substantially outperforms traditional non-pre-trained methods, significantly enhancing RM performance. For instance, POLAR-7B could improve preference accuracy from 54.8% to 81.0% on STEM tasks and from 57.9% to 85.5% on creative writing tasks compared to SOTA baselines. POLAR also shows robust generalization capabilities in RLHF using Reinforcement Fine-tuning (RFT), providing reliable reward signals and markedly enhancing policy performance—improving LLaMa3.1-8B from an average of 47.36% to 56.33% and Qwen2.5-32B from 64.49% to 70.47% on 20 benchmarks. Moreover, scaling experiments reveal a clear power-law relationship between computation and performance, supported by linear correlation coefficients approaching 0.99. The impressive performance, strong generalization, and scaling properties suggest that POLAR is a promising direction for developing general and strong reward models.
Shihan Dou, Shichun Liu, Yuming Yang 0001, Yicheng Zou, Yunhua Zhou, Shuhao Xing, Chenhao Huang, Qiming Ge, Haijun Lv, Demin Song, Songyang Gao, Chengqi Lyu, Enyu Zhou, Honglin Guo, Zhiheng Xi, Qipeng Guo, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001, Kai Chen 0026
NeurIPS14
2025 BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
abstract
In this paper, we introduce BMMR, a large-scale bilingual, multimodal, multi-disciplinary reasoning dataset for the community to develop and evaluate large multimodal models (LMMs). BMMR comprises 100k university-level questions drawn from 300 UNESCO-defined subjects, spanning diverse formats—multiple-choice, fill-in-the-blank, and open-ended QA—and sourced from both print and digital media such as books, exams, and quizzes. All data are curated and filtered via a human-in-the-loop, automated, and scalable framework, and each instance is paired with a high-quality reasoning path. The dataset is organized into two parts: BMMR-Eval that comprises 20k high-quality instances to comprehensively assess LMMs’ knowledge and reasoning across multiple disciplines in both Chinese and English; and BMMR-Train that contains 80k instances to support further research and development, extending the current focus on mathematical reasoning to diverse disciplines and domains. In addition, we propose the process-based multi-discipline BMMR-Verifier for accurate and fine-grained evaluation of LMMs’ reasoning. Extensive experiments reveal that (i) even SOTA models leave substantial headroom on BMMR-Eval; (ii) reasoning models exhibit discipline bias and outperform LMMs only on specific subjects; (iii) open-source models still trail their proprietary counterparts; and (iv) fine-tuning on BMMR-Train narrows this gap. Additionally, we conduct reasoning-chain analyses using BMMR-Verifier and other in-depth studies, uncovering the challenges LMMs currently face in multidisciplinary reasoning. We will release the data and models, and we believe our work can offers valuable insights and contributions to the community.
Zhiheng Xi, Yutao Fan, Honglin Guo, Yufang Liu, Xiaoran Fan, Jingchao Ding, Wangmeng Zuo, Zhenfei Yin, Lei Bai 0001, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
NeurIPS4
2025 CompCraft: Foreground-Driven Image Synthesis With Customized Layouts
abstract
Recently, advancements in text-to-image synthesis and image customization have drawn significant attention. Among these technologies, foreground-driven image synthesis models aim to create diverse scenes for specific foregrounds, showing broad application prospects. However, existing foreground-driven diffusion models struggle to accurately generate scenes with layouts that align with user intentions. To address these challenges, we propose CompCraft, a training-free framework that enhances layout control and improves overall generation quality in current models. First, CompCraft identifies that the failure of existing methods to achieve effective control arises from the excessive influence of fully denoised foreground information on the generated scene. To address this, we propose a foreground regularization strategy that modifies the foreground-related attention maps, reducing their impact and ensuring better integration of the foreground with the generated scene. Then, we propose a series of inference-time layout guidance strategies to guide the image generation process with the user’s finely customized layouts. These strategies enable current foreground-driven diffusion models with accurate layout control. Finally, we introduce a comprehensive benchmark to evaluate CompCraft. Both quantitative and qualitative results demonstrate that CompCraft can effectively generate high-quality images with precise customized layouts, showcasing its strong capabilities in pratical image synthesis applications.
Honglin Guo, Rui-dong Chen, Weizhi Nie, Lanjun Wang, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.1
2024 Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning
abstract
In this paper, we propose R$^3$: Learning Reasoning through Reverse Curriculum Reinforcement Learning (RL), a novel method that employs only outcome supervision to achieve the benefits of process supervision for large language models. The core challenge in applying RL to complex reasoning is to identify a sequence of actions that result in positive rewards and provide appropriate supervision for optimization. Outcome supervision provides sparse rewards for final results without identifying error locations, whereas process supervision offers step-wise rewards but requires extensive manual annotation. R$^3$ overcomes these limitations by learning from correct demonstrations. Specifically, R$^3$ progressively slides the start state of reasoning from a demonstration’s end to its beginning, facilitating easier model exploration at all stages. Thus, R$^3$ establishes a step-wise curriculum, allowing outcome supervision to offer step-level signals and precisely pinpoint errors. Using Llama2-7B, our method surpasses RL baseline on eight reasoning tasks by $4.1$ points on average. Notably, in program-based reasoning, 7B-scale models perform comparably to larger models or closed-source models with our R$^3$.
Zhiheng Xi, Wenxiang Chen, Boyang Hong, Senjie Jin, Wei He 0024, Yiwen Ding, Shichun Liu, Junzhe Wang 0001, Honglin Guo, Xiaoran Fan, Yuhao Zhou 0005, Shihan Dou, Xiao Wang 0001, Xinbo Zhang, Peng Sun 0006, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ICML11
2021 Fine-grained image inpainting with scale-enhanced generative adversarial network
Weirong Liu 0002, Chengrui Cao, Chenwen Ren, Yulin Wei, Honglin Guo
Pattern Recognit. Lett.6