Jiaxin Wen

dblp:189/3085 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 7 first-author · 10 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 CodePlan: Unlocking Reasoning Potential in Large Language Models by Scaling Code-form Planning
abstract
Despite the remarkable success of large language models (LLMs) on traditional natural language processing tasks, their planning ability remains a critical bottleneck in tackling complex multi-step reasoning tasks. Existing approaches mainly rely on prompting or task-specific fine-tuning, often suffering from weak robustness and cross-task generalization. To address the limitation, we introduce CodePlan, a scalable paradigm that empowers LLMs to generate and follow code-form plans---pseudocode that outlines high-level, structured reasoning processes. By leveraging the structured and versatile nature of code, CodePlan effectively captures the rich semantics and control flows inherent to sophisticated reasoning. Importantly, CodePlan allows the automatic extraction of code-form plans from massive, wide-ranging text corpora without the need for curated, task-specific datasets. This enables it to scale up efficiently and improve reasoning capabilities across diverse scenarios. To train CodePlan, we construct a large-scale dataset of 2M examples that integrate code-form plans with standard prompt-response pairs from existing corpora. With minimal computation overhead during both training and inference, CodePlan achieves a 25.1\% relative improvement compared with directly generating responses, averaged across 13 challenging multi-step reasoning benchmarks, spanning mathematical reasoning, symbolic reasoning, instruction-following, multi-hop QA, and decision-making tasks. Further analysis reveals CodePlan's increasing performance gains on more complex reasoning tasks, as well as significant data efficiency thanks to its generalization ability.
Jiaxin Wen, Jian Guan 0002, Hongning Wang, Wei Wu 0014, Minlie Huang
ICLR1
2025 Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
abstract
As large language models (LLMs) grow more powerful, they also become more difficult to trust. They could be either aligned with human intentions, or exhibit "subversive misalignment" -- introducing subtle errors that bypass safety checks. Although individual errors may not immediately cause harm, each increases the risk of an eventual safety failure. With this uncertainty, model deployment often grapples with the tradeoff between ensuring safety and harnessing the capabilities of untrusted models. In this work, we introduce the ``Diffuse Risk Management'' problem, aiming to balance the average-case safety and usefulness in the deployment of untrusted models over a large sequence of tasks. We approach this problem by developing a two-level framework: the single-task level (micro-protocol) and the whole-scenario level (macro-protocol). At the single-task level, we develop various \textit{micro}-protocols that use a less capable, but extensively tested (trusted) model to harness and monitor the untrusted model. At the whole-scenario level, we find an optimal \textit{macro}-protocol that uses an adaptive estimate of the untrusted model's risk to choose between micro-protocols. To evaluate the robustness of our method, we follow \textit{control evaluations} in a code generation testbed, which involves a red team attempting to generate subtly backdoored code with an LLM whose deployment is safeguarded by a blue team. Experiment results show that our approach retains 99.6\% usefulness of the untrusted model while ensuring near-perfect safety, significantly outperforming existing deployment methods. Our approach also demonstrates robustness when the trusted and untrusted models have a large capability gap. Our findings demonstrate the promise of managing diffuse risks in the deployment of increasingly capable but untrusted LLMs.
Jiaxin Wen, Vivek Hebbar, Caleb Larson, Aryan Bhatt, Ansh Radhakrishnan, Mrinank Sharma, Henry Sleight, Shi Feng 0005, He He 0001, Ethan Perez, Buck Shlegeris, Akbir Khan
ICLR1
2025 Language Models Learn to Mislead Humans via RLHF
abstract
Language models (LMs) can produce errors that are hard to detect for humans, especially when the task is complex. RLHF, the most popular post-training method, may exacerbate this problem: to achieve higher rewards, LMs might get better at convincing humans that they are right even when they are wrong. We study this phenomenon under a standard RLHF pipeline, calling it ``U-Sophistry'' since it is \textbf{U}nintended by model developers. Specifically, we ask time-constrained (e.g., 3-10 minutes) human subjects to evaluate the correctness of model outputs and calculate humans' accuracy against gold labels. On a question-answering task (QuALITY) and programming task (APPS), RLHF makes LMs better at convincing our subjects but not at completing the task correctly. RLHF also makes the model harder to evaluate: our subjects' false positive rate increases by 24.1% on QuALITY and 18.3% on APPS. Finally, we show that probing, a state-of-the-art approach for detecting \textbf{I}ntended Sophistry (e.g.~backdoored LMs), does not generalize to U-Sophistry. Our results highlight an important failure mode of RLHF and call for more research in assisting humans to align them.
Jiaxin Wen, Ruiqi Zhong, Akbir Khan, Ethan Perez, Jacob Steinhardt, Minlie Huang, Samuel R. Bowman, He He 0001, Shi Feng 0005
ICLR1
2025 Predicting Empirical AI Research Outcomes with Language Models
abstract
Many promising-looking ideas in AI research fail to deliver, but their validation takes substantial human labor and compute. Predicting an idea's chance of success is thus crucial for accelerating empirical AI research, a skill that even expert researchers can only acquire through substantial experience. We build the first benchmark for this task and compare LMs with human experts. Concretely, given two research ideas (e.g., two jailbreaking methods), we aim to predict which will perform better on a set of benchmarks. We scrape ideas and experimental results from conference papers, yielding 1,585 human-verified idea pairs \textit{published after our base model's cut-off date} for testing, and 6,000 pairs for training. We then develop a system that combines a fine-tuned GPT-4.1 with a paper retrieval agent, and we recruit 25 human experts to compare with. In the NLP domain, our system beats human experts by a large margin (64.4\% v.s. 48.9\%). On the full test set, our system achieves 77\% accuracy, while off-the-shelf frontier LMs like o3 perform no better than random guessing, even with the same retrieval augmentation. We verify that our system does not exploit superficial features like idea complexity through extensive human-written and LM-designed robustness tests. Finally, we evaluate our system on unpublished novel ideas, including ideas generated by an AI ideation agent. Our system achieves 63.6\% accuracy, demonstrating its potential as a reward model for improving idea generation models. Altogether, our results outline a promising new direction for LMs to accelerate empirical AI research.
Jiaxin Wen, Chenglei Si, Yueh-Han Chen, He He 0001, Shi Feng 0005
NeurIPS1
2024 Learning Task Decomposition to Assist Humans in Competitive Programming
abstract
When using language models (LMs) to solve complex problems, humans might struggle to understand the LM-generated solutions and repair the flawed ones.To assist humans in repairing them, we propose to automatically decompose complex solutions into multiple simpler pieces that correspond to specific subtasks.We introduce a novel objective for learning task decomposition, termed assistive value (AssistV), which measures the feasibility and speed for humans to repair the decomposed solution.We collect a dataset of human repair experiences on different decomposed solutions.Utilizing the collected data as in-context examples, we then learn to critique, refine, and rank decomposed solutions to improve AssistV.We validate our method under competitive programming problems: under 177 hours of human study, our method enables non-experts to solve 33.3% more problems, speeds them up by 3.3x, and empowers them to match unassisted experts.
Jiaxin Wen, Ruiqi Zhong, Pei Ke, Zhihong Shao, Hongning Wang, Minlie Huang
ACL (1)1
2024 Nefis: A network coding based flexible device-to-device video streaming scheme
Jun Yin 0004, Jiaxin Wen, Ming Zhu 0017, Lei Wang 0054
J. Netw. Comput. Appl.2
2023 ETHICIST: Targeted Training Data Extraction Through Loss Smoothed Soft Prompting and Calibrated Confidence Estimation
abstract
Large pre-trained language models achieve impressive results across many tasks.However, recent works point out that pre-trained language models may memorize a considerable fraction of their training data, leading to the privacy risk of information leakage.In this paper, we propose a method named ETHICIST for targeted training data Extraction THrough loss smoothed soft prompting and calIbrated ConfIdence eSTimation, investigating how to recover the suffix in the training data when given a prefix.To elicit memorization in the attacked model, we tune soft prompt embeddings while keeping the model fixed.We further propose a smoothing loss that smooths the loss distribution of the suffix tokens to make it easier to sample the correct suffix.In order to select the most probable suffix from a collection of sampled suffixes and estimate the prediction confidence, we propose a calibrated confidence estimation method, which normalizes the confidence of the generated suffixes with a local estimation.We show that ETHICIST significantly improves the extraction performance on a recently proposed public benchmark.We also investigate several factors influencing the data extraction performance, including decoding strategy, model scale, prefix length, and suffix length.Our code is available at https://github.com/ thu-coai/
Zhexin Zhang, Jiaxin Wen, Minlie Huang
ACL (1)2
2023 Unveiling the Implicit Toxicity in Large Language Models
abstract
The open-endedness of large language models (LLMs) combined with their impressive capabilities may lead to new safety issues when being exploited for malicious use.While recent studies primarily focus on probing toxic outputs that can be easily detected with existing toxicity classifiers, we show that LLMs can generate diverse implicit toxic outputs that are exceptionally difficult to detect via simply zero-shot prompting.Moreover, we propose a reinforcement learning (RL) based attacking method to further induce the implicit toxicity in LLMs.Specifically, we optimize the language model with a reward that prefers implicit toxic outputs to explicit toxic and non-toxic ones.Experiments on five widely-adopted toxicity classifiers demonstrate that the attack success rate can be significantly improved through RL fine-tuning.For instance, the RL-finetuned LLaMA-13B model achieves an attack success rate of 90.04% on BAD and 62.85% on Davinci003.Our findings suggest that LLMs pose a significant threat in generating undetectable implicit toxic outputs.We further show that fine-tuning toxicity classifiers on the annotated examples from our attacking method can effectively enhance their ability to detect LLM-generated implicit toxic language.The code is publicly available at https://github. com/thu-coai/Implicit-Toxicity.
Jiaxin Wen, Pei Ke, Hao Sun 0012, Zhexin Zhang, Chengfei Li, Jinfeng Bai, Minlie Huang
EMNLP1
2023 Re³Dial: Retrieve, Reorganize and Rescale Conversations for Long-Turn Open-Domain Dialogue Pre-training
abstract
Pre-training on large-scale open-domain dialogue data can substantially improve the performance of dialogue models.However, the pre-trained dialogue model's ability to utilize long-range context is limited due to the scarcity of long-turn dialogue sessions.Most dialogues in existing pre-training corpora contain fewer than three turns of dialogue.To alleviate this issue, we propose the Retrieve, Reorganize and Rescale framework (Re 3 Dial), which can automatically construct billion-scale long-turn dialogues by reorganizing existing short-turn ones.Given a short-turn session, Re 3 Dial first employs a session retriever to retrieve coherent consecutive sessions.To this end, we train the retriever to capture semantic and discourse relations within multi-turn dialogues through contrastive training.Next, Re 3 Dial samples a session from retrieved results following a diversity sampling strategy, which is designed to penalize repetitive or generic sessions.A longer session is then derived by concatenating the original session and the sampled session.By repeating the above process, Re 3 Dial can yield a coherent long-turn dialogue.Extensive experiments on multiple multi-turn dialogue benchmarks demonstrate that Re 3 Dial significantly improves the dialogue model's ability to utilize long-range context and thus generate more sensible and informative responses.Finally, we build a toolkit for efficiently rescaling conversations with Re 3 Dial, which enables us to construct a corpus containing 1B Chinese dialogue sessions with 11.3 turns on average (5× longer than the original corpus).Our retriever model, code, and data is publicly available at https://github.com/thu-coai/Re3Dial.
Jiaxin Wen, Hao Zhou 0012, Jian Guan 0002, Jie Zhou 0016, Minlie Huang
EMNLP1
2022 Persona-Guided Planning for Controlling the Protagonist's Persona in Story Generation
abstract
Endowing the protagonist with a specific personality is essential for writing an engaging story.In this paper, we aim to control the protagonist's persona in story generation, i.e., generating a story from a leading context and a persona description, where the protagonist should exhibit the specified personality through a coherent event sequence.Considering that personas are usually embodied implicitly and sparsely in stories, we propose a planning-based generation model named CONPER to explicitly model the relationship between personas and events.CON-PER first plans events of the protagonist's behavior which are motivated by the specified persona through predicting one target sentence, then plans the plot as a sequence of keywords with the guidance of the predicted persona-related events and commonsense knowledge, and finally generates the whole story.Both automatic and manual evaluation results demonstrate that CONPER outperforms state-of-the-art baselines for generating more coherent and persona-controllable stories.Our code is available at https:// github.com/thu-coai/ConPer.
Zhexin Zhang, Jiaxin Wen, Jian Guan 0002, Minlie Huang
NAACL-HLT2
2021 Robustness Testing of Language Understanding in Task-Oriented Dialog
abstract
Jiexi Liu, Ryuichi Takanobu, Jiaxin Wen, Dazhen Wan, Hongguang Li, Weiran Nie, Cheng Li, Wei Peng, Minlie Huang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jiexi Liu 0002, Ryuichi Takanobu, Jiaxin Wen, Dazhen Wan, Weiran Nie, Cheng Li 0040, Wei Peng 0011, Minlie Huang
ACL/IJCNLP (1)3
2016 Down-scaling SRTM slope based on histogram matching and slope distribution model
abstract
Terrain parameters are often extracted by down-scaling a lower resolution DEM at regional scale. The 1" SRTM (Shuttle Radar Topography Mission) elevation data is used widely at regional scale in terrain analysis. Slope is one of the most important terrain parameters. However, compared with the higher resolution slope surface, slope based on 1" SRTM elevation data reduced because of the relative lower resolution. Slope extracted from 1" SRTM was down-scaled in Shaanxi province of China in this paper in order to improve the data utility. The method of histogram matching and the theory of slope distribution model were used to develop a slope down-scaling model. After down-scaling in a reference area, 1" SRTM slope showed 95% similarity with the higher resolution slope, which improved 22% comparing with 1" SRTM slope before down-scaling. The slope data based on 1" SRTM, with better quality was obtained and is available for public use.
Qinke Yang, Jiaxin Wen, David L. B. Jupp
IGARSS3