Rui Wang 0092

dblp:06/2293-92 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Mitigating Context Interference for Reliable and Efficient Search Agents
abstract
Boyang Xue, Bin Wu, Shuofei Qiao, Sheng Wang, Rui Wang, Yiming Du, Hongru Wang, Jeff Z. Pan, Emine Yilmaz, Kam-Fai Wong, Aldo Lipani. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Boyang Xue, Bin Wu 0025, Shuofei Qiao, Rui Wang 0092, Yiming Du, Hongru Wang 0003, Jeff Z. Pan, Emine Yilmaz, Kam-Fai Wong, Aldo Lipani
ACL (1)5
2026 Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors
abstract
Zhiwei Zhang, Fei Zhao, Rui Wang, Zezhong Wang, Bin Liang, Jiakang Wang, Yao Hu, Shaosheng Cao, Kam-Fai Wong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Fei Zhao 0012, Rui Wang 0092, Zezhong Wang 0004, Bin Liang 0004, Jiakang Wang, Yao Hu 0002, Shaosheng Cao, Kam-Fai Wong
ACL (1)3
2026 EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
abstract
Abstract The pursuit of artificial general intelligence (AGI) has been accelerated by Multimodal Large Language Models (MLLMs), which exhibit superior reasoning, generalization capabilities, and proficiency in processing multimodal inputs. A crucial milestone in the evolution of AGI is the attainment of human-level planning, a fundamental ability for making informed decisions in complex environments, and solving a wide range of real-world problems. Despite the impressive advancements in MLLMs, a question remains: How far are current MLLMs from achieving human-level planning? To shed light on this question, we introduce EgoPlan-Bench, a comprehensive benchmark to evaluate the planning abilities of MLLMs in real-world scenarios from an egocentric perspective, mirroring human perception. EgoPlan-Bench emphasizes the evaluation of planning capabilities of MLLMs, featuring realistic tasks, diverse action plans, and intricate visual observations. Our rigorous evaluation of a wide range of MLLMs reveals that EgoPlan-Bench poses significant challenges, highlighting a substantial scope for improvement in MLLMs to achieve human-level task planning. To facilitate this advancement, we further present EgoPlan-IT, a specialized instruction-tuning dataset that effectively enhances model performance on EgoPlan-Bench. We have made all the codes, data, and a maintained benchmark leaderboard available at https://chenyi99.github.io/ego_plan/ to advance future research.
Yi Chen 0019, Yuying Ge, Yixiao Ge, Mingyu Ding, Bohao Li 0002, Rui Wang 0092, Ruifeng Xu 0001, Ying Shan, Xihui Liu
Int. J. Comput. Vis.6
2025 UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language Models
abstract
Boyang Xue, Fei Mi, Qi Zhu, Hongru Wang, Rui Wang, Sheng Wang, Erxin Yu, Xuming Hu, Kam-Fai Wong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Boyang Xue, Fei Mi, Qi Zhu 0007, Hongru Wang 0003, Rui Wang 0092, Erxin Yu, Xuming Hu, Kam-Fai Wong
ACL (1)5
2025 FGVIrony: A Chinese Dataset of Fine-grained Verbal Irony
Rui Wang 0092, Qianlong Wang 0001, Lin Gui 0003, Bin Liang 0004, Min Yang 0007, Ruifeng Xu 0001
Inf. Process. Manag.2
2024 UniRetriever: Multi-task Candidates Selection for Various Context-Adaptive Conversational Retrieval
abstract
Conversational retrieval refers to an information retrieval system that operates in an iterative and interactive manner, requiring the retrieval of various external resources, such as persona, knowledge, and even response, to effectively engage with the user and successfully complete the dialogue. However, most previous work trained independent retrievers for each specific resource, resulting in sub-optimal performance and low efficiency. Thus, we propose a multi-task framework function as a universal retriever for three dominant retrieval tasks during the conversation: persona selection, knowledge selection, and response selection. To this end, we design a dual-encoder architecture consisting of a context-adaptive dialogue encoder and a candidate encoder, aiming to attention to the relevant context from the long dialogue and retrieve suitable candidates by simply a dot product. Furthermore, we introduce two loss constraints to capture the subtle relationship between dialogue context and different candidates by regarding historically selected candidates as hard negatives. Extensive experiments and analysis establish state-of-the-art retrieval quality both within and outside its training domain, revealing the promising potential and generalization capability of our model to serve as a universal retriever for different candidate selection tasks simultaneously.
Hongru Wang 0003, Boyang Xue, Baohang Zhou, Rui Wang 0092, Fei Mi, Weichao Wang, Yasheng Wang, Kam-Fai Wong
LREC/COLING4
2024 SEED-Bench: Benchmarking Multimodal Large Language Models
abstract
Multimodal large language models (MLLMs), building upon the foundation of powerful large language models (LLMs), have recently demonstrated exceptional capabilities in generating not only texts but also images given in-terleaved multimodal inputs (acting like a combination of GPT-4V and DALL-E 3). However, existing MLLM benchmarks remain limited to assessing only models' comprehension ability of single image-text inputs, failing to keep up with the strides made in MLLMs. A comprehensive benchmark is imperative for investigating the progress and uncovering the limitations of current MLLMs. In this work, we categorize the capabilities of MLLMs into hierarchical levels from L0to L4based on the modalities they can ac-cept and generate, and propose SEED-Bench, a comprehensive benchmark that evaluates the hierarchical capa-bilities of MLLMs. Specifically, SEED-Bench comprises 24K multiple-choice questions with accurate human annotations, which span 27 dimensions, including the evaluation of both text and image generation. Multiple-choice questions with ground truth options derived from human annotation enable an objective and efficient assessment of model performance, eliminating the need for human or GPT intervention during evaluation. We further evaluate the performance of 22 prominent open-source MLLMs and summarize valuable observations. By revealing the limitations of existing MLLMs through extensive evaluations, we aim for SEED-Bench to provide insights that will mo-tivate future research toward the goal of General Artificial Intelligence. Dataset and evaluation code are available at https://github.com/AILab-CVC/SEED-Bench.
Bohao Li 0002, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang 0092, Ruimao Zhang, Ying Shan
CVPR5
2024 AppBench: Planning of Multiple APIs from Various APPs for Complex User Instruction
abstract
Large Language Models (LLMs) can interact with the real world by connecting with versatile external APIs, resulting in better problemsolving and task automation capabilities.Previous research primarily focuses on APIs with limited arguments from a single source or overlooks the complex dependency relationship between different APIs.However, it is essential to utilize multiple APIs collaboratively from various sources (e.g., different Apps in the iPhone), especially for complex user instructions.In this paper, we introduce AppBench, the first benchmark to evaluate LLMs' ability to plan and execute multiple APIs from various sources in order to complete the user's task.Specifically, we consider two significant challenges in multiple APIs: 1) graph structures: some APIs can be executed independently while others need to be executed one by one, resulting in graph-like execution order; and 2) permission constraints: which source is authorized to execute the API call.We have experimental results on 9 distinct LLMs; e.g., GPT-4o achieves only a 2.0% success rate at the most complex instruction, revealing that the existing state-of-the-art LLMs still cannot perform well in this situation even with the help of in-context learning and finetuning.
Hongru Wang 0003, Rui Wang 0092, Boyang Xue, Heming Xia, Jingtao Cao, Zeming Liu, Jeff Z. Pan, Kam-Fai Wong
EMNLP2
2024 Enhancing Large Language Models Against Inductive Instructions with Dual-critique Prompting
abstract
Rui Wang, Hongru Wang, Fei Mi, Boyang Xue, Yi Chen, Kam-Fai Wong, Ruifeng Xu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Rui Wang 0092, Hongru Wang 0003, Fei Mi, Boyang Xue, Yi Chen 0007, Kam-Fai Wong, Ruifeng Xu 0001
NAACL-HLT1
2024 TPE: Towards Better Compositional Reasoning over Cognitive Tools via Multi-persona Collaboration
Hongru Wang 0003, Lingzhi Wang 0001, Minda Hu, Rui Wang 0092, Boyang Xue, Kam-Fai Wong
NLPCC (2)5
2023 A Synthetic Data Generation Framework for Grounded Dialogues
abstract
Training grounded response generation models often requires a large collection of grounded dialogues.However, it is costly to build such dialogues.In this paper, we present a synthetic data generation framework (SynDG) for grounded dialogues.The generation process utilizes large pre-trained language models and freely available knowledge data (e.g., Wikipedia pages, persona profiles, etc.).The key idea of designing SynDG is to consider dialogue flow and coherence in the generation process.Specifically, given knowledge data, we first heuristically determine a dialogue flow, which is a series of knowledge pieces.Then, we employ T5 to incrementally turn the dialogue flow into a dialogue.To ensure coherence of both the dialogue flow and the synthetic dialogue, we design a two-level filtering strategy, at the flow-level and the utterance-level respectively.Experiments on two public benchmarks show that the synthetic grounded dialogue data produced by our framework is able to significantly boost model performance in both full training data and low-resource scenarios.
Jianzhu Bao, Rui Wang 0092, Yasheng Wang, Aixin Sun, Fei Mi, Ruifeng Xu 0001
ACL (1)2
2023 Retrieval-free Knowledge Injection through Multi-Document Traversal for Dialogue Models
abstract
Rui Wang, Jianzhu Bao, Fei Mi, Yi Chen, Hongru Wang, Yasheng Wang, Yitong Li, Lifeng Shang, Kam-Fai Wong, Ruifeng Xu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Rui Wang 0092, Jianzhu Bao, Fei Mi, Yi Chen 0007, Hongru Wang 0003, Yasheng Wang, Lifeng Shang, Kam-Fai Wong, Ruifeng Xu 0001
ACL (1)1
2022 Masking and Generation: An Unsupervised Method for Sarcasm Detection
abstract
Existing approaches for sarcasm detection are mainly based on supervised learning, in which the promising performance largely depends on a considerable amount of labeled data or extra information. In the real world scenario, however, the abundant labeled data or extra information requires high labor cost, not to mention that sufficient annotated data is unavailable in many low-resource conditions. To alleviate this dilemma, we investigate sarcasm detection from an unsupervised perspective, in which we explore a masking and generation paradigm in the context to extract the context incongruities for learning sarcastic expression. Further, to improve the feature representations of the sentences, we use unsupervised contrastive learning to improve the sentence representation based on the standard dropout. Experimental results on six perceived sarcasm detection benchmark datasets show that our approach outperforms baselines. Simultaneously, our unsupervised method obtains comparative performance with supervised methods for the intended sarcasm dataset.
Rui Wang 0092, Qianlong Wang 0001, Bin Liang 0004, Yi Chen 0019, Bing Qin 0001, Ruifeng Xu 0001
SIGIR1
2021 A Hierarchical Sequence Labeling Model for Argument Pair Extraction
Qinglin Zhu, Jianzhu Bao, Jipeng Wu, Caihua Yang, Rui Wang 0092, Ruifeng Xu 0001
NLPCC (2)6