Rujun Han

dblp:228/5627 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
14since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 9 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context
abstract
Long-horizon tasks that require sustained reasoning and multiple tool interactions remain challenging for LLM agents: small errors compound across steps, and even state-of-the-art models often hallucinate or lose coherence.We identify context management as the central bottleneck-extended histories cause agents to overlook critical evidence or become distracted by irrelevant information, thus failing to replan or reflect from previous mistakes.To address this, we propose COMPASS (Context-Organized Multi-Agent Planning and Strategy System), a lightweight hierarchical framework that separates tactical execution, strategic oversight, and context organization into three specialized components: (1) a Main Agent that performs reasoning and tool use, (2) a Meta-Thinker that monitors progress and issues strategic interventions, and (3) a Context Manager that maintains concise, relevant progress briefs for different reasoning stages.Across three challenging benchmarks-GAIA, BrowseComp, and Humanity's Last Exam-COMPASS improves accuracy by up to 20% relative to both single-and multi-agent baselines.We further introduce a test-time scaling extension that elevates performance to match established DeepResearch agents, and a posttraining pipeline that delegates context management to smaller models for enhanced efficiency.
Guangya Wan, Mingyang Ling 0001, Xiaoqi Ren, Rujun Han, Sheng Li 0001
ACL (1)4
2025 In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents
abstract
Large Language Models (LLMs) have made significant progress in open-ended dialogue, yet their inability to retain and retrieve relevant information from long-term interactions limits their effectiveness in applications requiring sustained personalization. External memory mechanisms have been proposed to address this limitation, enabling LLMs to maintain conversational continuity. However, existing approaches struggle with two key challenges. First, rigid memory granularity fails to capture the natural semantic structure of conversations, leading to fragmented and incomplete representations. Second, fixed retrieval mechanisms cannot adapt to diverse dialogue contexts and user interaction patterns. In this work, we propose Reflective Memory Management (RMM), a novel mechanism for long-term dialogue agents, integrating forward- and backward-looking reflections: (1) Prospective Reflection, which dynamically summarizes interactions across granularities—utterances, turns, and sessions—into a personalized memory bank for effective future retrieval, and (2) Retrospective Reflection, which iteratively refines the retrieval in an online reinforcement learning (RL) manner based on LLMs’ cited evidence. Experiments show that RMM demonstrates consistent improvement across various metrics and benchmarks. For example, RMM shows more than 10% accuracy improvement over the baseline without memory management on the LongMemEval dataset.
Zhen Tan 0001, Jun Yan 0001, I-Hung Hsu, Rujun Han, Zifeng Wang 0002, Long T. Le, Yiwen Song, Yanfei Chen, Hamid Palangi, Anand Rajan Iyer, Tianlong Chen 0001, Huan Liu 0001, Chen-Yu Lee, Tomas Pfister
ACL (1)4
2025 CiteEval: Principle-Driven Citation Evaluation for Source Attribution
abstract
Yumo Xu, Peng Qi, Jifan Chen, Kunlun Liu, Rujun Han, Lan Liu, Bonan Min, Vittorio Castelli, Arshit Gupta, Zhiguo Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yumo Xu, Peng Qi 0003, Jifan Chen, Kunlun Liu, Rujun Han, Lan Liu 0004, Bonan Min, Vittorio Castelli, Arshit Gupta, Zhiguo Wang 0006
ACL (1)5
2025 Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling
abstract
Recent advances in knowledge distillation (KD) have enabled smaller student models to approach the performance of larger teacher models. However, popular methods such as supervised KD and on-policy KD, are adversely impacted by the knowledge gaps between teacher-student in practical scenarios. Supervised KD suffers from a distribution mismatch between training with a static dataset and inference over final student-generated outputs. Conversely, on-policy KD, which uses student-generated samples for training, can suffer from low-quality training examples with which teacher models are not familiar, resulting in inaccurate teacher feedback. To address these limitations, we introduce Speculative Knowledge Distillation (SKD), a novel approach that leverages cooperation between student and teacher models to generate high-quality training data on-the-fly while aligning with the student's inference-time distribution. In SKD, the student proposes tokens, and the teacher replaces poorly ranked ones based on its own distribution, transferring high-quality knowledge adaptively. We evaluate SKD on various text generation tasks, including translation, summarization, math, and instruction following, and show that SKD consistently outperforms existing KD methods across different domains, data sizes, and model initialization strategies.
Wenda Xu, Rujun Han, Zifeng Wang 0002, Long T. Le, Dhruv Madeka, Lei Li 0005, William Yang Wang, Rishabh Agarwal, Chen-Yu Lee, Tomas Pfister
ICLR2
2025 Reverse Thinking Makes LLMs Stronger Reasoners
abstract
Justin Chen, Zifeng Wang, Hamid Palangi, Rujun Han, Sayna Ebrahimi, Long Le, Vincent Perot, Swaroop Mishra, Mohit Bansal, Chen-Yu Lee, Tomas Pfister. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Justin Chih-Yao Chen, Zifeng Wang 0002, Hamid Palangi, Rujun Han, Sayna Ebrahimi, Long T. Le, Vincent Perot, Swaroop Mishra, Mohit Bansal, Chen-Yu Lee, Tomas Pfister
NAACL (Long Papers)4
2024 RAG-QA Arena: Evaluating Domain Robustness for Long-form Retrieval Augmented Question Answering
abstract
Rujun Han, Yuhao Zhang, Peng Qi, Yumo Xu, Jenyuan Wang, Lan Liu, William Yang Wang, Bonan Min, Vittorio Castelli. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Rujun Han, Yuhao Zhang 0004, Peng Qi 0003, Yumo Xu, Jenyuan Wang, Lan Liu 0004, William Yang Wang, Bonan Min, Vittorio Castelli
EMNLP1
2024 Dancing in Chains: Reconciling Instruction Following and Faithfulness in Language Models
abstract
Modern language models (LMs) need to follow human instructions while being faithful; yet, they often fail to achieve both.Here, we provide concrete evidence of a trade-off between instruction following (i.e., follow open-ended instructions) and faithfulness (i.e., ground responses in given context) when training LMs with these objectives.For instance, fine-tuning LLaMA-7B on instruction following datasets renders it less faithful.Conversely, instructiontuned Vicuna-7B shows degraded performance at following instructions when further optimized on tasks that require contextual grounding.One common remedy is multi-task learning (MTL) with data mixing, yet it remains far from achieving a synergic outcome.We propose a simple yet effective method that relies on Rejection Sampling for Continued Selfinstruction Tuning (RESET), which significantly outperforms vanilla MTL.Surprisingly, we find that less is more, as training RESET with high-quality, yet substantially smaller data (three-fold less) yields superior results.Our findings offer a better understanding of objective discrepancies in alignment training of LMs.
Zhengxuan Wu, Yuhao Zhang 0004, Peng Qi 0003, Yumo Xu, Rujun Han, Yian Zhang, Jifan Chen, Bonan Min, Zhiheng Huang
EMNLP5
2023 ACCENT: An Automatic Event Commonsense Evaluation Metric for Open-Domain Dialogue Systems
abstract
Commonsense reasoning is omnipresent in human communications and thus is an important feature for open-domain dialogue systems.However, evaluating commonsense in dialogue systems is still an open challenge.We take the first step by focusing on event commonsense that considers events and their relations, and is crucial in both dialogues and general commonsense reasoning.We propose AC-CENT, an event commonsense evaluation metric empowered by commonsense knowledge bases (CSKBs).ACCENT first extracts eventrelation tuples from a dialogue, and then evaluates the response by scoring the tuples in terms of their compatibility with the CSKB.To evaluate ACCENT, we construct the first public event commonsense evaluation dataset for open-domain dialogues.Our experiments show that ACCENT is an efficient metric for event commonsense evaluation, which achieves higher correlations with human judgments than existing baselines.
Sarik Ghazarian, Yijia Shao, Rujun Han, Aram Galstyan, Nanyun Peng 0001
ACL (1)3
2022 Character-centric Story Visualization via Visual Planning and Token Alignment
abstract
Story visualization advances the traditional text-to-image generation by enabling multiple image generation based on a complete story.This task requires machines to 1) understand long text inputs and 2) produce a globally consistent image sequence that illustrates the contents of the story.A key challenge of consistent story visualization is to preserve characters that are essential in stories.To tackle the challenge, we propose to adapt a recent work that augments Vector-Quantized Variational Autoencoders (VQ-VAE) with a text-tovisual-token (transformer) architecture.Specifically, we modify the text-to-visual-token module with a two-stage framework: 1) character token planning model that predicts the visual tokens for characters only; 2) visual token completion model that generates the remaining visual token sequence, which is sent to VQ-VAE for finalizing image generations.To encourage characters to appear in the images, we further train the two-stage framework with a character-token alignment objective.Extensive experiments and evaluations demonstrate that the proposed method excels at preserving characters and can produce higher quality image sequences compared with the strong baselines.Code can be found in https: //github.com/PlusLabNLP/VP-CSV
Hong Chen 0017, Rujun Han, Te-Lin Wu, Hideki Nakayama, Nanyun Peng 0001
EMNLP2
2022 Go Back in Time: Generating Flashbacks in Stories with Event Temporal Prompts
abstract
Stories or narratives are comprised of a sequence of events.To compose interesting stories, professional writers often leverage a creative writing technique called flashback that inserts past events into current storylines as we commonly observe in novels and plays.However, it is challenging for machines to generate flashbacks as it requires solid understanding of event temporal order (e.g.feeling hungry before eat, not vice versa), and the creativity to arrange storylines so that earlier events do not always appear first in narrative order.Two major issues in existing systems exacerbate the challenges: 1) temporal bias in pretraining and story datasets that leads to monotonic event temporal orders; 2) lack of explicit guidance that helps machines decide where to insert flashbacks.We propose to address these issues using structured storylines to encode events and their pair-wise temporal relations ( before , after and vague ) as temporal prompts that guide how stories should unfold temporally.We leverage a Plan-and-Write framework enhanced by reinforcement learning to generate storylines and stories end-toend.Evaluation results show that the proposed method can generate more interesting stories with flashbacks while maintaining textual diversity, fluency and temporal coherence.1
Rujun Han, Hong Chen 0017, Yufei Tian, Nanyun Peng 0001
NAACL-HLT1
2021 Clinical Temporal Relation Extraction with Probabilistic Soft Logic Regularization and Global Inference
abstract
There has been a steady need in the medical community to precisely extract the temporal relations between clinical events. In particular, temporal information can facilitate a variety of downstream applications such as case report retrieval and medical question answering. Existing methods either require expensive feature engineering or are incapable of modeling the global relational dependencies among the events. In this paper, we propose a novel method, Clinical Temporal ReLation Exaction with Probabilistic Soft Logic Regularization and Global Inference (CTRL-PG) to tackle the problem at the document level. Extensive experiments on two benchmark datasets, I2B2-2012 and TB-Dense, demonstrate that CTRL-PG significantly outperforms baseline methods for temporal relation extraction.
Yichao Zhou 0001, Rujun Han, J. Harry Caufield, Kai-Wei Chang 0001, Yizhou Sun, Peipei Ping, Wei Wang 0010
AAAI3
2021 Modeling Context in Answer Sentence Selection Systems on a Latency Budget
abstract
Answer Sentence Selection (AS2) is an efficient approach for the design of open-domain Question Answering (QA) systems.In order to achieve low latency, traditional AS2 models score question-answer pairs individually, ignoring any information from the document each potential answer was extracted from.In contrast, more computationally expensive models designed for machine reading comprehension tasks typically receive one or more passages as input, which often results in better accuracy.In this work, we present an approach to efficiently incorporate contextual information in AS2 models.For each answer candidate, we first use unsupervised similarity techniques to extract relevant sentences from its source document, which we then feed into an efficient transformer architecture fine-tuned for AS2.Our best approach, which leverages a multi-way attention architecture to efficiently encode context, improves 6% to 11% over noncontextual state of the art in AS2 with minimal impact on system latency.All experiments in this work were conducted in English.* Work was conducted while the author was an intern at Amazon Alexa.The math of pi explained, as simply as possible How many digits of pi we really need?Pi, you may also remember from grade school, is not an ordinary number.It's irrational, meaning it has an endless number of decimals that never repeat.Though even cutting off pi at 15 digits allows for extremely precise measurements.If you were to draw a circle with a diameter of 25 billion miles, using 15 digits of pi, you'd only arrive at a measurement of the circumference that's off by 1.5 inches, NASA's Marc Rayman explained in a post on NASA's JPL website.And that's good enough.Of course, that hasn't stopped people from looking for more and more digits of pi.Currently, there are more than 22.4 trillion known digits, which show no hint of ending or repeating.Further reading: pi and pie.
Rujun Han, Luca Soldaini, Alessandro Moschitti
EACL1
2021 ECONET: Effective Continual Pretraining of Language Models for Event Temporal Reasoning
abstract
While pre-trained language models (PTLMs) have achieved noticeable success on many NLP tasks, they still struggle for tasks that require event temporal reasoning, which is essential for event-centric applications.We present a continual pre-training approach that equips PTLMs with targeted knowledge about event temporal relations.We design self-supervised learning objectives to recover masked-out event and temporal indicators and to discriminate sentences from their corrupted counterparts (where event or temporal indicators got replaced).By further pre-training a PTLM with these objectives jointly, we reinforce its attention to event and temporal information, yielding enhanced capability on event temporal reasoning.This Effective CONtinual pre-training framework for Event Temporal reasoning (ECONET) improves the PTLMs' fine-tuning performances across five relation extraction and question answering tasks and achieves new or on-par state-of-the-art performances in most of our downstream tasks. 1
Rujun Han, Xiang Ren 0001, Nanyun Peng 0001
EMNLP (1)1
2021 ESTER: A Machine Reading Comprehension Dataset for Reasoning about Event Semantic Relations
abstract
Understanding how events are semantically related to each other is the essence of reading comprehension.Recent event-centric reading comprehension datasets focus mostly on event arguments or temporal relations.While these tasks partially evaluate machines' ability of narrative understanding, human-like reading comprehension requires the capability to process event-based information beyond arguments and temporal reasoning.For example, to understand causality between events, we need to infer motivation or purpose; to establish event hierarchy, we need to understand the composition of events.To facilitate these tasks, we introduce ESTER, a comprehensive machine reading comprehension (MRC) dataset for Event Semantic Relation Reasoning.The dataset leverages natural language queries to reason about the five most common event semantic relations, provides more than 6K questions, and captures 10.1K event relation pairs.Experimental results show that the current SOTA systems achieve 22.1%, 63.3% and 83.5% for token-based exact-match (EM), F 1 and event-based HIT@1 scores, which are all significantly below human performances (36.0%, 79.6%, 100% respectively), highlighting our dataset as a challenging benchmark.1
Rujun Han, I-Hung Hsu, Jiao Sun, Julia Baylon, Qiang Ning, Dan Roth 0001, Nanyun Peng 0001
EMNLP (1)1
2020 Domain Knowledge Empowered Structured Neural Net for End-to-End Event Temporal Relation Extraction
abstract
Extracting event temporal relations is a critical task for information extraction and plays an important role in natural language understanding.Prior systems leverage deep learning and pre-trained language models to improve the performance of the task.However, these systems often suffer from two shortcomings: 1) when performing maximum a posteriori (MAP) inference based on neural models, previous systems only used structured knowledge that is assumed to be absolutely correct, i.e., hard constraints; 2) biased predictions on dominant temporal relations when training with a limited amount of data.To address these issues, we propose a framework that enhances deep neural network with distributional constraints constructed by probabilistic domain knowledge.We solve the constrained inference problem via Lagrangian Relaxation and apply it to end-to-end event temporal relation extraction tasks.Experimental results show our framework is able to improve the baseline neural network models with strong statistical significance on two widely used datasets in news and clinical domains.
Rujun Han, Yichao Zhou 0001, Nanyun Peng 0001
EMNLP (1)1
2020 TORQUE: A Reading Comprehension Dataset of Temporal Ordering Questions
abstract
A critical part of reading is being able to understand the temporal relationships between events described in a passage of text, even when those relationships are not explicitly stated.However, current machine reading comprehension benchmarks have practically no questions that test temporal phenomena, so systems trained on these benchmarks have no capacity to answer questions such as "what happened before/after [some event]?"We introduce TORQUE, a new English reading comprehension benchmark built on 3.2k news snippets with 21k human-generated questions querying temporal relationships.Results show that RoBERTa-large achieves an exact-match score of 51% on the test set of TORQUE, about 30% behind human performance.1 1 https://allennlp.org/torque.htmlHeavy snow is causing disruption to transport across the UK, with heavy rainfall bringing flooding to the south-west of England.Rescuers searching for a woman trapped in a landslide at her home in Looe, Cornwall, said they had found a body.Q1: What events have already finished?A: searching trapped landslide said found Q2: What events have begun but has not finished?A: snow causing disruption rainfall bringing flooding Q3: What will happen in the future?A: No answers.Q4: What happened before a woman was trapped?A: landslide Q5: What had started before a woman was trapped?A: snow rainfall landslide Q6: What happened while a woman was trapped?A: searching Q7: What happened after a woman was trapped?A: searching said found Q8: What happened at about the same time as the snow?A: rainfall Q9: What happened after the snow started?A: causing disruption bringing flooding searching trapped landslide said found Q10: What happened before the snow started?A: No answers.warm
Qiang Ning, Hao Wu 0034, Rujun Han, Nanyun Peng 0001, Matt Gardner 0001, Dan Roth 0001
EMNLP (1)3
2019 Deep Structured Neural Network for Event Temporal Relation Extraction
abstract
We propose a novel deep structured learning framework for event temporal relation extraction.The model consists of 1) a recurrent neural network (RNN) to learn scoring functions for pair-wise relations, and 2) a structured support vector machine (SSVM) to make joint predictions.The neural network automatically learns representations that account for long-term contexts to provide robust features for the structured model, while the SSVM incorporates domain knowledge such as transitive closure of temporal relations as constraints to make better globally consistent decisions.By jointly training the two components, our model combines the benefits of both data-driven learning and knowledge exploitation.Experimental results on three highquality event temporal relation datasets (TCR, MATRES, and TB-Dense) demonstrate that incorporated with pre-trained contextualized embeddings, the proposed model achieves significantly better performances than the stateof-the-art methods on all three datasets.We also provide thorough ablation studies to investigate our model.
Rujun Han, I-Hung Hsu, Mu Yang, Aram Galstyan, Ralph M. Weischedel, Nanyun Peng 0001
CoNLL1
2019 Joint Event and Temporal Relation Extraction with Shared Representations and Structured Prediction
abstract
Rujun Han, Qiang Ning, Nanyun Peng. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Rujun Han, Qiang Ning, Nanyun Peng 0001
EMNLP/IJCNLP (1)1
2018 Conditional Word Embedding and Hypothesis Testing via Bayes-by-Backprop
abstract
Conventional word embedding models do not leverage information from document metadata, and they do not model uncertainty.We address these concerns with a model that incorporates document covariates to estimate conditional word embedding distributions.Our model allows for (a) hypothesis tests about the meanings of terms, (b) assessments as to whether a word is near or far from another conditioned on different covariate values, and (c) assessments as to whether estimated differences are statistically significant.
Rujun Han, Michael Gill, Arthur Spirling, Kyunghyun Cho
EMNLP1