Hongru Wang 0003

dblp:72/1462-3 · DBLP profile ↗
← Back
29ranked-venue papers
9as first author
29since 2021 · last 2026
0000-0001-5027-0138ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 6 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Mem²Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation
abstract
Zihao Cheng, Zeming Liu, Yingyu Shan, Xinyi Wang, Xiangrong Zhu, Yunpu Ma, Hongru Wang, Yuhang Guo, Wei Lin, Yunhong Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zeming Liu, Yingyu Shan, Xiangrong Zhu 0002, Yunpu Ma, Hongru Wang 0003, Yuhang Guo 0001, Yunhong Wang 0001
ACL (1)7
2026 From Word to World: Can Large Language Models be Implicit Text-based World Models?
abstract
Yixia Li, Hongru Wang, Jiahao Qiu, Zhenfei Yin, Dongdong Zhang, Cheng Qian, Zeping Li, Xiaoteng Ma, Guanhua Chen, Heng Ji. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yixia Li, Hongru Wang 0003, Jiahao Qiu, Zhenfei Yin, Dongdong Zhang 0001, Cheng Qian 0008, Zeping Li, Xiaoteng Ma, Guanhua Chen 0001, Heng Ji 0001
ACL (1)2
2026 Rethinking the Role of Entropy in Optimizing Tool-Use Behaviors for Large Language Model Agents
abstract
Zeping Li, Hongru Wang, Yiwen Zhao, Guanhua Chen, Yixia Li, Keyang Chen, Yixin Cao, Guangnan Ye, Hongfeng Chai, Zhenfei Yin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zeping Li, Hongru Wang 0003, Guanhua Chen 0001, Yixia Li, Keyang Chen, Yixin Cao 0002, Guangnan Ye, Hongfeng Chai, Zhenfei Yin
ACL (1)2
2026 WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation Models
abstract
Rui Wang, Ce Zhang, Jun-Yu Ma, Jianshu Zhang, Hongru Wang, Yi Chen, Boyang Xue, Tianqing Fang, Zhisong Zhang, Hongming Zhang, Haitao Mi, Dong Yu, Kam-Fai Wong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Rui Wang 0015, Ce Zhang 0009, Jun-Yu Ma, Hongru Wang 0003, Yi Chen 0007, Boyang Xue, Tianqing Fang, Zhisong Zhang, Hongming Zhang 0009, Haitao Mi, Dong Yu 0001, Kam-Fai Wong
ACL (1)5
2026 Illusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood Consistency
abstract
Haoming Xu, Ningyuan Zhao, Yunzhi Yao, Weihong Xu, Hongru Wang, Xinle Deng, Shumin Deng, Jeff Z. Pan, Huajun Chen, Ningyu Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ningyuan Zhao, Yunzhi Yao, Hongru Wang 0003, Xinle Deng, Shumin Deng, Jeff Z. Pan, Huajun Chen, Ningyu Zhang 0001
ACL (1)5
2026 Mitigating Context Interference for Reliable and Efficient Search Agents
abstract
Boyang Xue, Bin Wu, Shuofei Qiao, Sheng Wang, Rui Wang, Yiming Du, Hongru Wang, Jeff Z. Pan, Emine Yilmaz, Kam-Fai Wong, Aldo Lipani. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Boyang Xue, Bin Wu 0025, Shuofei Qiao, Rui Wang 0092, Yiming Du, Hongru Wang 0003, Jeff Z. Pan, Emine Yilmaz, Kam-Fai Wong, Aldo Lipani
ACL (1)7
2025 Can LLMs Evaluate Complex Attribution in QA? Automatic Benchmarking using Knowledge Graphs
abstract
Nan Hu, Jiaoyan Chen, Yike Wu, Guilin Qi, Hongru Wang, Sheng Bi, Yongrui Chen, Tongtong Wu, Jeff Z. Pan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Nan Hu 0004, Jiaoyan Chen 0001, Guilin Qi, Hongru Wang 0003, Yongrui Chen 0002, Tongtong Wu, Jeff Z. Pan
ACL (1)5
2025 UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language Models
abstract
Boyang Xue, Fei Mi, Qi Zhu, Hongru Wang, Rui Wang, Sheng Wang, Erxin Yu, Xuming Hu, Kam-Fai Wong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Boyang Xue, Fei Mi, Qi Zhu 0007, Hongru Wang 0003, Rui Wang 0092, Erxin Yu, Xuming Hu, Kam-Fai Wong
ACL (1)4
2025 ImPart: Importance-Aware Delta-Sparsification for Improved Model Compression and Merging in LLMs
abstract
With the proliferation of task-specific large language models, delta compression has emerged as a method to mitigate the resource challenges of deploying numerous such models by effectively compressing the delta model parameters. Previous delta-sparsification methods either remove parameters randomly or truncate singular vectors directly after singular value decomposition (SVD). However, these methods either disregard parameter importance entirely or evaluate it with too coarse a granularity. In this work, we introduce ImPart, a novel importance-aware delta sparsification approach. Leveraging SVD, it dynamically adjusts sparsity ratios of different singular vectors based on their importance, effectively retaining crucial task-specific knowledge even at high sparsity ratios. Experiments show that ImPart achieves state-of-the-art delta sparsification performance, demonstrating 2\times higher compression ratio than baselines at the same performance level. When integrated with existing methods, ImPart sets a new state-of-the-art on delta quantization and model merging.
Yixia Li, Hongru Wang 0003, Xuetao Wei, James Jian Qiao Yu, Yun Chen 0007, Guanhua Chen 0001
ACL (1)3
2025 NILE: Internal Consistency Alignment in Large Language Models
abstract
Minda Hu, Qiyuan Zhang, Yufei Wang, Bowei He, Hongru Wang, Jingyan Zhou, Liangyou Li, Yasheng Wang, Chen Ma, Irwin King. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Minda Hu, Qiyuan Zhang 0001, Yufei Wang 0005, Bowei He, Hongru Wang 0003, Jingyan Zhou, Liangyou Li, Yasheng Wang, Chen Ma 0001, Irwin King
EMNLP5
2025 Self-DC: When to Reason and When to Act? Self Divide-and-Conquer for Compositional Unknown Questions
abstract
Hongru Wang, Boyang Xue, Baohang Zhou, Tianhua Zhang, Cunxiang Wang, Huimin Wang, Guanhua Chen, Kam-Fai Wong. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Hongru Wang 0003, Boyang Xue, Baohang Zhou, Tianhua Zhang, Cunxiang Wang, Guanhua Chen 0001, Kam-Fai Wong
NAACL (Long Papers)1
2025 SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters
abstract
Yan Yang, Zeguan Xiao, Xin Lu, Hongru Wang, Xuetao Wei, Hailiang Huang, Guanhua Chen, Yun Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Zeguan Xiao, Hongru Wang 0003, Xuetao Wei, Guanhua Chen 0001, Yun Chen 0007
NAACL (Long Papers)4
2025 Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering
abstract
Yu Zhao, Alessio Devoto, Giwon Hong, Xiaotang Du, Aryo Pradipta Gema, Hongru Wang, Xuanli He, Kam-Fai Wong, Pasquale Minervini. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yu Zhao 0043, Alessio Devoto, Giwon Hong, Xiaotang Du, Aryo Pradipta Gema, Hongru Wang 0003, Xuanli He, Kam-Fai Wong, Pasquale Minervini
NAACL (Long Papers)6
2024 JoTR: A Joint Transformer and Reinforcement Learning Framework for Dialogue Policy Learning
abstract
Dialogue policy learning (DPL) aims to determine an abstract representation (also known as action) to guide what the response should be. Typically, DPL is cast as a sequential decision problem across a series of predefined action candidates. However, such static and narrow actions can limit response diversity and impede the dialogue agent’s adaptability to new scenarios and edge cases. To overcome these challenges, we introduce a novel Joint Transformer Reinforcement Learning framework, coined as JoTR, where a text-to-text Transformer-based model is employed to directly generate dialogue actions. More concretely, JoTR formulates a token-grained policy, facilitating more dynamic and adaptable dialogue action generation without the need for predefined action candidates. This method not only enhances the diversity of responses but also significantly improves the system’s capability to manage unfamiliar scenarios. Furthermore, JoTR utilizes Reinforcement Learning with a reward-shaping mechanism to efficiently fine-tune the token-grained policy. This allows the model to evolve through interactions, thereby enhancing its performance over time. Our extensive evaluation demonstrates that JoTR surpasses previous state-of-the-art models, showing improvements of 9% and 13% in success rate, and 34% and 37% in the diversity of dialogue actions across two benchmark dialogue modeling tasks respectively. These results have been validated by both user simulators and human evaluators. Code and data are available at ://github.com/KwanWaiChung/JoTR.
Wai-Chung Kwan, Hongru Wang 0003, Zezhong Wang 0004, Bin Liang 0004, Xian Wu 0001, Yefeng Zheng 0001, Kam-Fai Wong
LREC/COLING3
2024 UniRetriever: Multi-task Candidates Selection for Various Context-Adaptive Conversational Retrieval
abstract
Conversational retrieval refers to an information retrieval system that operates in an iterative and interactive manner, requiring the retrieval of various external resources, such as persona, knowledge, and even response, to effectively engage with the user and successfully complete the dialogue. However, most previous work trained independent retrievers for each specific resource, resulting in sub-optimal performance and low efficiency. Thus, we propose a multi-task framework function as a universal retriever for three dominant retrieval tasks during the conversation: persona selection, knowledge selection, and response selection. To this end, we design a dual-encoder architecture consisting of a context-adaptive dialogue encoder and a candidate encoder, aiming to attention to the relevant context from the long dialogue and retrieve suitable candidates by simply a dot product. Furthermore, we introduce two loss constraints to capture the subtle relationship between dialogue context and different candidates by regarding historically selected candidates as hard negatives. Extensive experiments and analysis establish state-of-the-art retrieval quality both within and outside its training domain, revealing the promising potential and generalization capability of our model to serve as a universal retriever for different candidate selection tasks simultaneously.
Hongru Wang 0003, Boyang Xue, Baohang Zhou, Rui Wang 0092, Fei Mi, Weichao Wang, Yasheng Wang, Kam-Fai Wong
LREC/COLING1
2024 MCIL: Multimodal Counterfactual Instance Learning for Low-resource Entity-based Multimodal Information Extraction
abstract
Multimodal information extraction (MIE) is a challenging task which aims to extract the structural information in free text coupled with the image for constructing the multimodal knowledge graph. The entity-based MIE tasks are based on the entity information to complete the specific tasks. However, the existing methods only investigated the entity-based MIE tasks under supervised learning with adequate labeled data. In the real-world scenario, collecting enough data and annotating the entity-based samples are time-consuming, and impractical. Therefore, we propose to investigate the entity-based MIE tasks under the low-resource settings. The conventional models are prone to overfitting on limited labeled data, which can result in poor performance. This is because the models tend to learn the bias existing in the limited samples, which can lead them to model the spurious correlations between multimodal features and task labels. To provide a more comprehensive understanding of the bias inherent in multimodal features of MIE samples, we decompose the features into image, entity, and context factors. Furthermore, we investigate the causal relationships between these factors and model performance, leveraging the structural causal model to delve into the correlations between the input features and output labels. Based on this, we propose the multimodal counterfactual instance learning framework to generate the counterfactual instances by the interventions on the limited observational samples. In the framework, we analyze the causal effect of the counterfactual instances and exploit it as a supervisory signal to maximize the effect for reducing the bias and improving the generalization of the model. Empirically, we evaluate the proposed method on the two public MIE benchmark datasets and the experimental results verify the effectiveness of it.
Baohang Zhou, Ying Zhang 0015, Kehui Song, Hongru Wang 0003, Yu Zhao 0043, Xuhui Sui, Xiaojie Yuan
LREC/COLING4
2024 AppBench: Planning of Multiple APIs from Various APPs for Complex User Instruction
abstract
Large Language Models (LLMs) can interact with the real world by connecting with versatile external APIs, resulting in better problemsolving and task automation capabilities.Previous research primarily focuses on APIs with limited arguments from a single source or overlooks the complex dependency relationship between different APIs.However, it is essential to utilize multiple APIs collaboratively from various sources (e.g., different Apps in the iPhone), especially for complex user instructions.In this paper, we introduce AppBench, the first benchmark to evaluate LLMs' ability to plan and execute multiple APIs from various sources in order to complete the user's task.Specifically, we consider two significant challenges in multiple APIs: 1) graph structures: some APIs can be executed independently while others need to be executed one by one, resulting in graph-like execution order; and 2) permission constraints: which source is authorized to execute the API call.We have experimental results on 9 distinct LLMs; e.g., GPT-4o achieves only a 2.0% success rate at the most complex instruction, revealing that the existing state-of-the-art LLMs still cannot perform well in this situation even with the help of in-context learning and finetuning.
Hongru Wang 0003, Rui Wang 0092, Boyang Xue, Heming Xia, Jingtao Cao, Zeming Liu, Jeff Z. Pan, Kam-Fai Wong
EMNLP1
2024 VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models
abstract
Progress in Text-to-Image (T2I) models has significantly advanced the generation of images from textual descriptions.Existing metrics, such as CLIP, effectively measure the semantic alignment between single prompts and their corresponding images.However, they fall short in evaluating a model's ability to generalize across a broad spectrum of textual inputs.To address this gap, we propose the VLEU (Visual Language Evaluation Understudy) metric.VLEU leverages the power of Large Language Models (LLMs) to sample from the visual text domain, encompassing the entire range of potential inputs for the T2I task, to generate a wide variety of visual text.The images generated by T2I models from these prompts are then assessed for their alignment with the input text using the CLIP model.VLEU quantitatively measures a model's generalizability by computing the Kullback-Leibler (KL) divergence between the visual text marginal distribution and the conditional distribution over the images generated by the model.This provides a comprehensive metric for comparing the overall generalizability of T2I models, beyond single-prompt evaluations, and offers valuable insights during the finetuning process.Our experimental results demonstrate VLEU's effectiveness in evaluating the generalizability of various T2I models, positioning it as an essential metric for future research and development in image synthesis from text prompts.
Jingtao Cao, Zheng Zhang 0064, Hongru Wang 0003, Kam-Fai Wong
EMNLP3
2024 Knowledge Conflicts for LLMs: A Survey
abstract
This survey provides an in-depth analysis of knowledge conflicts for large language models (LLMs), highlighting the complex challenges they encounter when blending contextual and parametric knowledge.Our focus is on three categories of knowledge conflicts: contextmemory, inter-context, and intra-memory conflict.These conflicts can significantly impact the trustworthiness and performance of LLMs, especially in real-world applications where noise and misinformation are common.By categorizing these conflicts, exploring the causes, examining the behaviors of LLMs under such conflicts, and reviewing available solutions, this survey aims to shed light on strategies for improving the robustness of LLMs, thereby serving as a valuable resource for advancing research in this evolving area.
Rongwu Xu, Zehan Qi, Zhijiang Guo, Cunxiang Wang, Hongru Wang 0003, Yue Zhang 0004, Wei Xu 0039
EMNLP5
2024 M3sum: A Novel Unsupervised Language-Guided Video Summarization
abstract
Language-guided video summarization empowers users to use natural language queries to effortlessly summarize lengthy videos into concise and relevant summaries that cater specifically to their information needs, which is more friendly to access and digest. However, most of the previous works rely on tremendous (also expensive) annotated videos and complex designs to align different modals at the feature level. In this paper, we first explore the combination of off-the-shelf models for each modal to solve the complex multi-modal problem by proposing a novel unsupervised language-guided video summarization method: Modular Multi-Modal Summarization (M3Sum), which does not require any training data or parameter updates. Specifically, instead of training an alignment module at the feature level, we convert all modal information (e.g. audio and frames) into textual descriptions and design a parameter-free alignment mechanism to fuse text descriptions from different modals. Benefiting from the remarkable long-context understanding capability of large language models (LLMs), our approach demonstrates comparable performance to most unsupervised methods and even outperforms certain supervised methods.
Hongru Wang 0003, Baohang Zhou, Zhengkun Zhang, Yiming Du, David Ho, Kam-Fai Wong
ICASSP1
2024 SELF-GUARD: Empower the LLM to Safeguard Itself
abstract
Zezhong Wang, Fangkai Yang, Lu Wang, Pu Zhao, Hongru Wang, Liang Chen, Qingwei Lin, Kam-Fai Wong. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zezhong Wang 0004, Fangkai Yang, Lu Wang 0029, Pu Zhao 0004, Hongru Wang 0003, Liang Chen 0001, Qingwei Lin, Kam-Fai Wong
NAACL-HLT5
2024 Enhancing Large Language Models Against Inductive Instructions with Dual-critique Prompting
abstract
Rui Wang, Hongru Wang, Fei Mi, Boyang Xue, Yi Chen, Kam-Fai Wong, Ruifeng Xu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Rui Wang 0092, Hongru Wang 0003, Fei Mi, Boyang Xue, Yi Chen 0007, Kam-Fai Wong, Ruifeng Xu 0001
NAACL-HLT2
2024 AutoPSV: Automated Process-Supervised Verifier
abstract
In this work, we propose a novel method named \textbf{Auto}mated \textbf{P}rocess-\textbf{S}upervised \textbf{V}erifier (\textbf{\textsc{AutoPSV}}) to enhance the reasoning capabilities of large language models (LLMs) by automatically annotating the reasoning steps. \textsc{AutoPSV} begins by training a verification model on the correctness of final answers, enabling it to generate automatic process annotations. This verification model assigns a confidence score to each reasoning step, indicating the probability of arriving at the correct final answer from that point onward. We detect relative changes in the verification's confidence scores across reasoning steps to automatically annotate the reasoning process, enabling error detection even in scenarios where ground truth answers are unavailable. This alleviates the need for numerous manual annotations or the high computational costs associated with model-induced annotation approaches. We experimentally validate that the step-level confidence changes learned by the verification model trained on the final answer correctness can effectively identify errors in the reasoning steps. We demonstrate that the verification model, when trained on process annotations generated by \textsc{AutoPSV}, exhibits improved performance in selecting correct answers from multiple LLM-generated outputs. Notably, we achieve substantial improvements across five datasets in mathematics and commonsense reasoning. The source code of \textsc{AutoPSV} is available at \url{https://github.com/rookie-joe/AutoPSV}.
Jianqiao Lu, Zhiyang Dou, Hongru Wang 0003, Zeyu Cao, Jianbo Dai, Yunlong Feng, Zhijiang Guo
NeurIPS3
2024 TPE: Towards Better Compositional Reasoning over Cognitive Tools via Multi-persona Collaboration
Hongru Wang 0003, Lingzhi Wang 0001, Minda Hu, Rui Wang 0092, Boyang Xue, Kam-Fai Wong
NLPCC (2)1
2024 Empowering Large Language Models: Tool Learning for Real-World Interaction
abstract
Since the advent of large language models (LLMs), the field of tool learning has remained very active in solving various tasks in practice, including but not limited to information retrieval. This half-day tutorial provides basic concepts of this field and an overview of recent advancements with several applications. In specific, we start with some foundational components and architecture of tool learning (i.e., cognitive tool and physical tool), and then we categorize existing studies in this field into tool-augmented learning and tool-oriented learning, and introduce various learning methods to empower LLMs this kind of capability. Furthermore, we provide several cases about when, what, and how to use tools in different applications. We end with some open challenges and several potential research directions for future studies. We believe this tutorial is suited for both researchers at different stages (introductory, intermediate, and advanced) and industry practitioners who are interested in LLMs and tool learning.
Hongru Wang 0003, Yujia Qin, Yankai Lin 0001, Jeff Z. Pan, Kam-Fai Wong
SIGIR1
2024 KddRES: A Multi-level Knowledge-driven Dialogue Dataset for Restaurant Towards Customized Dialogue System
Hongru Wang 0003, Wai-Chung Kwan, Zimo Zhou, Kam-Fai Wong
Comput. Speech Lang.1
2023 Retrieval-free Knowledge Injection through Multi-Document Traversal for Dialogue Models
abstract
Rui Wang, Jianzhu Bao, Fei Mi, Yi Chen, Hongru Wang, Yasheng Wang, Yitong Li, Lifeng Shang, Kam-Fai Wong, Ruifeng Xu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Rui Wang 0092, Jianzhu Bao, Fei Mi, Yi Chen 0007, Hongru Wang 0003, Yasheng Wang, Lifeng Shang, Kam-Fai Wong, Ruifeng Xu 0001
ACL (1)5
2023 MCML: A Novel Memory-based Contrastive Meta-Learning Method for Few Shot Slot Tagging
abstract
Hongru Wang, Zezhong Wang, Wai Chung Kwan, Kam-Fai Wong. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Hongru Wang 0003, Zezhong Wang 0004, Wai Chung Kwan, Kam-Fai Wong
IJCNLP (1)1
2022 Integrating Pretrained Language Model for Dialogue Policy Evaluation
abstract
Reinforcement Learning (RL) has been witnessed its potential for training a dialogue policy agent towards maximizing the accumulated rewards given from users. However, the reward can be very sparse for it is usually only provided at the end of a dialog session, which causes unaffordable interaction requirements for an acceptable dialog agent. Distinguished from many efforts dedicated to optimizing the policy and recovering the reward alternatively which suffers from easily getting stuck in local optima and model collapse, we decompose the adversarial training into two steps: 1) we integrate a pre-trained language model as a discriminator to judge whether the current system action is good enough for the last user action (i.e., next action prediction); 2) the discriminator gives and extra local dense reward to guide the agent’s exploration. The experimental result demonstrates that our method significantly improves the complete rate (4.4%) and success rate ( 8.0%) of the dialogue system.
Hongru Wang 0003, Zezhong Wang 0004, Kam-Fai Wong
ICASSP1