EDBT 2026 Demo / reviewers in the wild / expert
Zhichen Dong
dblp:368/8608
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 55% Trustworthy machine learning · 24% Reinforcement learning · 12% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 80% Automated reasoning and model checking · 20% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Mathematical optimization › discrete optimization
mixed integer linear programming |
1.5 | 2 | 2024 | Towards Imitation Learning to Branch for MIP: A Hybrid Reinforcement Learning based Sample Augmentation Approach · ICLR 2024 L2P-MIP: Learning to Presolve for Mixed Integer Programming · ICLR 2024 |
Machine learning › Trustworthy machine learning › interpretability
representation probing |
0.9 | 1 | 2025 | Emergent Response Planning in LLMs · ICML 2025 |
Natural language and speech › Language models and text generation
alignment |
0.8 | 1 | 2024 | Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models · NeurIPS 2024 |
Machine learning › Reinforcement learning
imitation learning |
0.8 | 1 | 2024 | Towards Imitation Learning to Branch for MIP: A Hybrid Reinforcement Learning based Sample Augmentation Approach · ICLR 2024 |
Natural language and speech › Language models and text generation › alignment
inference-time alignment |
0.8 | 1 | 2024 | Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models · NeurIPS 2024 |
Natural language and speech › Language models and text generation
large language model safety |
0.8 | 1 | 2024 | Emulated Disalignment: Safety Alignment for Large Language Models May Backfire! · ACL (1) 2024 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › search control
learning to branch |
0.8 | 1 | 2024 | Towards Imitation Learning to Branch for MIP: A Hybrid Reinforcement Learning based Sample Augmentation Approach · ICLR 2024 |
Machine learning › Trustworthy machine learning › AI safety
safety alignment |
0.8 | 1 | 2024 | Emulated Disalignment: Safety Alignment for Large Language Models May Backfire! · ACL (1) 2024 |
Natural language and speech › Language models and text generation › alignment › scalable oversight
weak-to-strong generalization |
0.8 | 1 | 2024 | Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models · NeurIPS 2024 |
Mathematical optimization › integer programming
branch-and-bound |
0.8 | 1 | 2024 | Towards Imitation Learning to Branch for MIP: A Hybrid Reinforcement Learning based Sample Augmentation Approach · ICLR 2024 |
Automated reasoning and model checking › satisfiability › SAT solving
branching heuristic |
0.8 | 1 | 2024 | Towards Imitation Learning to Branch for MIP: A Hybrid Reinforcement Learning based Sample Augmentation Approach · ICLR 2024 |
Mathematical optimization
presolving |
0.8 | 1 | 2024 | L2P-MIP: Learning to Presolve for Mixed Integer Programming · ICLR 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.3 | 1 | 2025 | Emergent Response Planning in LLMs · ICML 2025 |
Natural language and speech › Language models and text generation
instruction following |
0.2 | 1 | 2024 | Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models · NeurIPS 2024 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.2 | 1 | 2024 | Towards Imitation Learning to Branch for MIP: A Hybrid Reinforcement Learning based Sample Augmentation Approach · ICLR 2024 |
Natural language and speech › Language models and text generation › alignment
preference alignment |
0.2 | 1 | 2024 | Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
sample augmentation · 1.5online reinforcement learning · 1.5offline reinforcement learning · 1.5imitation learning · 1.5probing · 0.9test-time guidance · 0.8supervised learning · 0.8preference optimization · 0.8log-probability difference · 0.8heuristics · 0.8greedy search · 0.8adversarial training · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Emergent Response Planning in LLMsabstractIn this work, we argue that large language models (LLMs), though trained to predict only the next token, exhibit emergent planning behaviors: $\textbf{their hidden representations encode future outputs beyond the next token}$. Through simple probing, we demonstrate that LLM prompt representations encode global attributes of their entire responses, including $\textit{structure attributes}$ (e.g., response length, reasoning steps), $\textit{content attributes}$ (e.g., character choices in storywriting, multiple-choice answers at the end of response), and $\textit{behavior attributes}$ (e.g., answer confidence, factual consistency). In addition to identifying response planning, we explore how it scales with model size across tasks and how it evolves during generation. The findings that LLMs plan ahead for the future in their hidden representations suggest potential applications for improving transparency and generation control. Zhichen Dong, Zhanhui Zhou, Zhixuan Liu, Chao Yang 0026, Chaochao Lu |
ICML | 1 |
| 2024 | Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!abstractZhanhui Zhou, Jie Liu, Zhichen Dong, Jiaheng Liu, Chao Yang, Wanli Ouyang, Yu Qiao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhanhui Zhou, Jie Liu 0047, Zhichen Dong, Chao Yang 0026, Wanli Ouyang, Yu Qiao 0001 |
ACL (1) | 3 |
| 2024 | L2P-MIP: Learning to Presolve for Mixed Integer ProgrammingabstractModern solvers for solving mixed integer programming (MIP) often rely on the branch-and-bound (B&B) algorithm which could be of high time complexity, and presolving techniques are well designed to simplify the instance as pre-processing before B&B. However, such presolvers in existing literature or open-source solvers are mostly set by default agnostic to specific input instances, and few studies have been reported on tailoring presolving settings. In this paper, we aim to dive into this open question and show that the MIP solver can be indeed largely improved when switching the default instance-agnostic presolving into instance-specific presolving. Specifically, we propose a combination of supervised learning and classic heuristics to achieve efficient presolving adjusting, avoiding tedious reinforcement learning. Notably, our approach is orthogonal from many recent efforts in incorporating learning modules into the B&B framework after the presolving stage, and to our best knowledge, this is the first work for introducing learning to presolve in MIP solvers. Experiments on multiple real-world datasets show that well-trained neural networks can infer proper presolving for arbitrary incoming MIP instances in less than 0.5s, which is neglectable compared with the solving time often hours or days. Chang Liu 0021, Zhichen Dong, Haobo Ma, Weilin Luo, Xijun Li, Junchi Yan |
ICLR | 2 |
| 2024 | Towards Imitation Learning to Branch for MIP: A Hybrid Reinforcement Learning based Sample Augmentation ApproachabstractBranch-and-bound (B\&B) has long been favored for tackling complex Mixed Integer Programming (MIP) problems, where the choice of branching strategy plays a pivotal role. Recently, Imitation Learning (IL)-based policies have emerged as potent alternatives to traditional rule-based approaches. However, it is nontrivial to acquire high-quality training samples, and IL often converges to suboptimal variable choices for branching, restricting the overall performance. In response to these challenges, we propose a novel hybrid online and offline reinforcement learning (RL) approach to enhance the branching policy by cost-effective training sample augmentation. In the online phase, we train an online RL agent to dynamically decide the sample generation processes, drawing from either the learning-based policy or the expert policy. The objective is to strike a balance between exploration and exploitation of the sample generation process. In the offline phase, a value function is trained to fit each decision's cumulative reward and filter the samples with high cumulative returns. This dual-purpose function not only reduces training complexity but also enhances the quality of the samples. To assess the efficacy of our data augmentation mechanism, we conduct comprehensive evaluations across a range of MIP problems. The results consistently show that it excels in making superior branching decisions compared to state-of-the-art learning-based models and the open-source solver SCIP. Notably, it even often outperforms Gurobi. Changwen Zhang, Wenli Ouyang, Hao Yuan 0002, Liming Gong, Ziao Guo, Zhichen Dong, Junchi Yan |
ICLR | 7 |
| 2024 | Attacks, Defenses and Evaluations for LLM Conversation Safety: A SurveyabstractZhichen Dong, Zhanhui Zhou, Chao Yang, Jing Shao, Yu Qiao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zhichen Dong, Zhanhui Zhou, Chao Yang 0026, Yu Qiao 0001 |
NAACL-HLT | 1 |
| 2024 | Weak-to-Strong Search: Align Large Language Models via Searching over Small Language ModelsabstractLarge language models are usually fine-tuned to align with human preferences. However, fine-tuning a large language model can be challenging. In this work, we introduce $\textit{weak-to-strong search}$, framing the alignment of a large language model as a test-time greedy search to maximize the log-probability difference between small tuned and untuned models while sampling from the frozen large model. This method serves both as (1) a compute-efficient model up-scaling strategy that avoids directly tuning the large model and as (2) an instance of weak-to-strong generalization that enhances a strong model with weak test-time guidance.
Empirically, we demonstrate the flexibility of weak-to-strong search across different tasks. In controlled-sentiment generation and summarization, we use tuned and untuned $\texttt{gpt2}$s to improve the alignment of large models without additional training. Crucially, in a more difficult instruction-following benchmark, AlpacaEval 2.0, we show that reusing off-the-shelf small models (e.g., $\texttt{zephyr-7b-beta}$ and its untuned version) can improve the length-controlled win rates of both white-box and black-box large models against $\texttt{gpt-4-turbo}$ (e.g., $34.4\% \rightarrow 37.9\%$ for $\texttt{Llama-3-70B-Instruct}$ and $16.0\% \rightarrow 20.1\%$ for $\texttt{gpt-3.5-turbo-instruct}$), despite the small models' low win rates $\approx 10.0\%$. Zhanhui Zhou, Zhixuan Liu, Jie Liu 0047, Zhichen Dong, Chao Yang 0026, Yu Qiao 0001 |
NeurIPS | 4 |