Zhichen Dong

dblp:368/8608 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 55% Trustworthy machine learning · 24% Reinforcement learning · 12%
Theoretical computer science
2 papers
Mathematical optimization · 80% Automated reasoning and model checking · 20%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Mathematical optimization › discrete optimization
mixed integer linear programming
1.522024
Towards Imitation Learning to Branch for MIP: A Hybrid Reinforcement Learning based Sample Augmentation Approach · ICLR 2024
L2P-MIP: Learning to Presolve for Mixed Integer Programming · ICLR 2024
Machine learning › Trustworthy machine learning › interpretability
representation probing
0.912025
Emergent Response Planning in LLMs · ICML 2025
Natural language and speech › Language models and text generation
alignment
0.812024
Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models · NeurIPS 2024
Machine learning › Reinforcement learning
imitation learning
0.812024
Towards Imitation Learning to Branch for MIP: A Hybrid Reinforcement Learning based Sample Augmentation Approach · ICLR 2024
Natural language and speech › Language models and text generation › alignment
inference-time alignment
0.812024
Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models · NeurIPS 2024
Natural language and speech › Language models and text generation
large language model safety
0.812024
Emulated Disalignment: Safety Alignment for Large Language Models May Backfire! · ACL (1) 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › search control
learning to branch
0.812024
Towards Imitation Learning to Branch for MIP: A Hybrid Reinforcement Learning based Sample Augmentation Approach · ICLR 2024
Machine learning › Trustworthy machine learning › AI safety
safety alignment
0.812024
Emulated Disalignment: Safety Alignment for Large Language Models May Backfire! · ACL (1) 2024
Natural language and speech › Language models and text generation › alignment › scalable oversight
weak-to-strong generalization
0.812024
Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models · NeurIPS 2024
Mathematical optimization › integer programming
branch-and-bound
0.812024
Towards Imitation Learning to Branch for MIP: A Hybrid Reinforcement Learning based Sample Augmentation Approach · ICLR 2024
Automated reasoning and model checking › satisfiability › SAT solving
branching heuristic
0.812024
Towards Imitation Learning to Branch for MIP: A Hybrid Reinforcement Learning based Sample Augmentation Approach · ICLR 2024
Mathematical optimization
presolving
0.812024
L2P-MIP: Learning to Presolve for Mixed Integer Programming · ICLR 2024
Machine learning › Trustworthy machine learning
interpretability
0.312025
Emergent Response Planning in LLMs · ICML 2025
Natural language and speech › Language models and text generation
instruction following
0.212024
Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models · NeurIPS 2024
Machine learning › Reinforcement learning
offline reinforcement learning
0.212024
Towards Imitation Learning to Branch for MIP: A Hybrid Reinforcement Learning based Sample Augmentation Approach · ICLR 2024
Natural language and speech › Language models and text generation › alignment
preference alignment
0.212024
Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

sample augmentation · 1.5online reinforcement learning · 1.5offline reinforcement learning · 1.5imitation learning · 1.5probing · 0.9test-time guidance · 0.8supervised learning · 0.8preference optimization · 0.8log-probability difference · 0.8heuristics · 0.8greedy search · 0.8adversarial training · 0.8
YearPublicationVenuePosition
2025 Emergent Response Planning in LLMs
abstract
In this work, we argue that large language models (LLMs), though trained to predict only the next token, exhibit emergent planning behaviors: $\textbf{their hidden representations encode future outputs beyond the next token}$. Through simple probing, we demonstrate that LLM prompt representations encode global attributes of their entire responses, including $\textit{structure attributes}$ (e.g., response length, reasoning steps), $\textit{content attributes}$ (e.g., character choices in storywriting, multiple-choice answers at the end of response), and $\textit{behavior attributes}$ (e.g., answer confidence, factual consistency). In addition to identifying response planning, we explore how it scales with model size across tasks and how it evolves during generation. The findings that LLMs plan ahead for the future in their hidden representations suggest potential applications for improving transparency and generation control.
Zhichen Dong, Zhanhui Zhou, Zhixuan Liu, Chao Yang 0026, Chaochao Lu
ICML1
2024 Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!
abstract
Zhanhui Zhou, Jie Liu, Zhichen Dong, Jiaheng Liu, Chao Yang, Wanli Ouyang, Yu Qiao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zhanhui Zhou, Jie Liu 0047, Zhichen Dong, Chao Yang 0026, Wanli Ouyang, Yu Qiao 0001
ACL (1)3
2024 L2P-MIP: Learning to Presolve for Mixed Integer Programming
abstract
Modern solvers for solving mixed integer programming (MIP) often rely on the branch-and-bound (B&B) algorithm which could be of high time complexity, and presolving techniques are well designed to simplify the instance as pre-processing before B&B. However, such presolvers in existing literature or open-source solvers are mostly set by default agnostic to specific input instances, and few studies have been reported on tailoring presolving settings. In this paper, we aim to dive into this open question and show that the MIP solver can be indeed largely improved when switching the default instance-agnostic presolving into instance-specific presolving. Specifically, we propose a combination of supervised learning and classic heuristics to achieve efficient presolving adjusting, avoiding tedious reinforcement learning. Notably, our approach is orthogonal from many recent efforts in incorporating learning modules into the B&B framework after the presolving stage, and to our best knowledge, this is the first work for introducing learning to presolve in MIP solvers. Experiments on multiple real-world datasets show that well-trained neural networks can infer proper presolving for arbitrary incoming MIP instances in less than 0.5s, which is neglectable compared with the solving time often hours or days.
Chang Liu 0021, Zhichen Dong, Haobo Ma, Weilin Luo, Xijun Li, Junchi Yan
ICLR2
2024 Towards Imitation Learning to Branch for MIP: A Hybrid Reinforcement Learning based Sample Augmentation Approach
abstract
Branch-and-bound (B\&B) has long been favored for tackling complex Mixed Integer Programming (MIP) problems, where the choice of branching strategy plays a pivotal role. Recently, Imitation Learning (IL)-based policies have emerged as potent alternatives to traditional rule-based approaches. However, it is nontrivial to acquire high-quality training samples, and IL often converges to suboptimal variable choices for branching, restricting the overall performance. In response to these challenges, we propose a novel hybrid online and offline reinforcement learning (RL) approach to enhance the branching policy by cost-effective training sample augmentation. In the online phase, we train an online RL agent to dynamically decide the sample generation processes, drawing from either the learning-based policy or the expert policy. The objective is to strike a balance between exploration and exploitation of the sample generation process. In the offline phase, a value function is trained to fit each decision's cumulative reward and filter the samples with high cumulative returns. This dual-purpose function not only reduces training complexity but also enhances the quality of the samples. To assess the efficacy of our data augmentation mechanism, we conduct comprehensive evaluations across a range of MIP problems. The results consistently show that it excels in making superior branching decisions compared to state-of-the-art learning-based models and the open-source solver SCIP. Notably, it even often outperforms Gurobi.
Changwen Zhang, Wenli Ouyang, Hao Yuan 0002, Liming Gong, Ziao Guo, Zhichen Dong, Junchi Yan
ICLR7
2024 Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
abstract
Zhichen Dong, Zhanhui Zhou, Chao Yang, Jing Shao, Yu Qiao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zhichen Dong, Zhanhui Zhou, Chao Yang 0026, Yu Qiao 0001
NAACL-HLT1
2024 Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models
abstract
Large language models are usually fine-tuned to align with human preferences. However, fine-tuning a large language model can be challenging. In this work, we introduce $\textit{weak-to-strong search}$, framing the alignment of a large language model as a test-time greedy search to maximize the log-probability difference between small tuned and untuned models while sampling from the frozen large model. This method serves both as (1) a compute-efficient model up-scaling strategy that avoids directly tuning the large model and as (2) an instance of weak-to-strong generalization that enhances a strong model with weak test-time guidance. Empirically, we demonstrate the flexibility of weak-to-strong search across different tasks. In controlled-sentiment generation and summarization, we use tuned and untuned $\texttt{gpt2}$s to improve the alignment of large models without additional training. Crucially, in a more difficult instruction-following benchmark, AlpacaEval 2.0, we show that reusing off-the-shelf small models (e.g., $\texttt{zephyr-7b-beta}$ and its untuned version) can improve the length-controlled win rates of both white-box and black-box large models against $\texttt{gpt-4-turbo}$ (e.g., $34.4\% \rightarrow 37.9\%$ for $\texttt{Llama-3-70B-Instruct}$ and $16.0\% \rightarrow 20.1\%$ for $\texttt{gpt-3.5-turbo-instruct}$), despite the small models' low win rates $\approx 10.0\%$.
Zhanhui Zhou, Zhixuan Liu, Jie Liu 0047, Zhichen Dong, Chao Yang 0026, Yu Qiao 0001
NeurIPS4