Haiyang Yu 0003

dblp:90/6643-3 · DBLP profile ↗
← Back
20ranked-venue papers
1as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 1 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
YearPublicationVenuePosition
2026 EvoRoute: Experience-Driven Self-Routing LLM Agent Systems
abstract
Guibin Zhang, Haiyang Yu, Kaiming Yang, Bingli Wu, Fei Huang, Yongbin Li, Shuicheng Yan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Guibin Zhang, Haiyang Yu 0003, Kaiming Yang, Bingli Wu, Fei Huang 0002, Yongbin Li 0001, Shuicheng Yan
ACL (1)2
2025 DeepSolution: Boosting Complex Engineering Solution Design via Tree-based Exploration and Bi-point Thinking
abstract
Zhuoqun Li, Haiyang Yu, Xuanang Chen, Hongyu Lin, Yaojie Lu, Fei Huang, Xianpei Han, Yongbin Li, Le Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Haiyang Yu 0003, Xuanang Chen, Yaojie Lu 0001, Fei Huang 0002, Xianpei Han, Yongbin Li 0001, Le Sun 0001
ACL (1)2
2025 IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference Optimization
abstract
In the realm of large language models (LLMs), the ability of models to accurately follow instructions is paramount as more agents and applications leverage LLMs for construction, where the complexity of instructions are rapidly increasing.However, on the one hand, there is only a certain amount of complex instruction evaluation data; on the other hand, there are no dedicated algorithms to improve the ability to follow complex instructions.To this end, this paper introduces TRACE, a benchmark for improving and evaluating the complex instructionfollowing ability, which consists of 120K training data and 1K evaluation data.Furthermore, we propose IOPO (Input-Output Preference Optimization) alignment method which takes both input and output preference pairs into consideration, where LLMs not only rapidly align with response preferences but also meticulously explore the instruction preferences.Extensive experiments on both in-domain and outof-domain datasets confirm the effectiveness of IOPO, showing 8.15%, 2.18% improvements on in-domain data and 5.91%, 2.83% on outof-domain data compared to SFT and DPO respectively.
Xinghua Zhang 0001, Haiyang Yu 0003, Cheng Fu 0003, Fei Huang 0002, Yongbin Li 0001
ACL (1)2
2025 EIFBENCH: Extremely Complex Instruction Following Benchmark for Large Language Models
abstract
With the development and widespread application of large language models (LLMs), the new paradigm of "Model as Product" is rapidly evolving, and demands higher capabilities to address complex user needs, often requiring precise workflow execution which involves the accurate understanding of multiple tasks.However, existing benchmarks focusing on single-task environments with limited constraints lack the complexity required to fully reflect real-world scenarios.To bridge this gap, we present the Extremely Complex Instruction Following Benchmark (EIFBENCH), meticulously crafted to facilitate a more realistic and robust evaluation of LLMs.EIFBENCH not only includes multi-task scenarios that enable comprehensive assessment across diverse task types concurrently, but also integrates a variety of constraints, replicating complex operational environments.Furthermore, we propose the Segment Policy Optimization (SegPO) algorithm to enhance the LLM's ability to accurately fulfill multi-task workflow.Evaluations on EIFBENCH have unveiled considerable performance discrepancies in existing LLMs when challenged with these extremely complex instructions.This finding underscores the necessity for ongoing optimization to navigate the intricate challenges posed by LLM applications.
Tao Zou 0003, Xinghua Zhang 0001, Haiyang Yu 0003, Minzheng Wang 0002, Fei Huang 0002, Yongbin Li 0001
EMNLP3
2025 StructRAG: Boosting Knowledge Intensive Reasoning of LLMs via Inference-time Hybrid Information Structurization
abstract
Retrieval-augmented generation (RAG) is a key means to effectively enhance large language models (LLMs) in many knowledge-based tasks. However, existing RAG methods struggle with knowledge-intensive reasoning tasks, because useful information required to these tasks are badly scattered. This characteristic makes it difficult for existing RAG methods to accurately identify key information and perform global reasoning with such noisy augmentation. In this paper, motivated by the cognitive theories that humans convert raw information into various structured knowledge when tackling knowledge-intensive reasoning, we proposes a new framework, StructRAG, which can identify the optimal structure type for the task at hand, reconstruct original documents into this structured format, and infer answers based on the resulting structure. Extensive experiments across various knowledge-intensive tasks show that StructRAG achieves state-of-the-art performance, particularly excelling in challenging scenarios, demonstrating its potential as an effective solution for enhancing LLMs in complex real-world applications.
Xuanang Chen, Haiyang Yu 0003, Yaojie Lu 0001, Qiaoyu Tang, Fei Huang 0002, Xianpei Han, Le Sun 0001, Yongbin Li 0001
ICLR3
2025 On the Role of Attention Heads in Large Language Model Safety
abstract
Large language models (LLMs) achieve state-of-the-art performance on multiple language tasks, yet their safety guardrails can be circumvented, leading to harmful generations. In light of this, recent research on safety mechanisms has emerged, revealing that when safety representations or component are suppressed, the safety capability of LLMs are compromised. However, existing research tends to overlook the safety impact of multi-head attention mechanisms, despite their crucial role in various model functionalities. Hence, in this paper, we aim to explore the connection between standard attention mechanisms and safety capability to fill this gap in the safety-related mechanistic interpretability. We propose an novel metric which tailored for multi-head attention, the Safety Head ImPortant Score (Ships), to assess the individual heads' contributions to model safety. Base on this, we generalize Ships to the dataset level and further introduce the Safety Attention Head AttRibution Algorithm (Sahara) to attribute the critical safety attention heads inside the model. Our findings show that special attention head has a significant impact on safety. Ablating a single safety head allows aligned model (e.g., Llama-2-7b-chat) to respond to **16$\times\uparrow$** more harmful queries, while only modifying **0.006\%** $\downarrow$ of the parameters, in contrast to the $\sim$ **5\%** modification required in previous studies. More importantly, we demonstrate that attention heads primarily function as feature extractors for safety and models fine-tuned from the same base model exhibit overlapping safety heads through comprehensive experiments. Together, our attribution approach and findings provide a novel perspective for unpacking the black box of safety mechanisms in large models.
Zhenhong Zhou, Haiyang Yu 0003, Xinghua Zhang 0001, Rongwu Xu, Fei Huang 0002, Kun Wang 0056, Yang Liu 0005, Junfeng Fang, Yongbin Li 0001
ICLR2
2025 Detecting AI-Generated Video via Frame Consistency
abstract
The increasing realism of AI-generated videos has raised potential security concerns, making it difficult for humans to distinguish them from the naked eye. Despite these concerns, limited research has been dedicated to detecting such videos effectively. To this end, we propose an open-source AI-generated video detection dataset. Our dataset spans diverse objects, scenes, behaviors, and actions by organizing input prompts into independent dimensions. It also includes various generation models with different generative models, featuring popular commercial models such as OpenAI’s Sora, Google’s Veo, and Kwai’s Kling. Furthermore, we propose a simple yet effective Detection model based on Concistency of Frame (DeCoF), which learns robust temporal artifacts across different generation methods. Extensive experiments demonstrate the generality and efficacy of the proposed DeCoF in detecting AI-generated videos, including those from nowadays’ mainstream commercial generators.
Qinglang Guo, Yong Liao 0003, Haiyang Yu 0003, Peng Yuan Zhou
ICME5
2025 Transferable Post-training via Inverse Value Learning
abstract
Xinyu Lu, Xueru Wen, Yaojie Lu, Bowen Yu, Hongyu Lin, Haiyang Yu, Le Sun, Xianpei Han, Yongbin Li. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Xueru Wen, Yaojie Lu 0001, Bowen Yu 0002, Haiyang Yu 0003, Le Sun 0001, Xianpei Han, Yongbin Li 0001
NAACL (Long Papers)6
2024 Preference Ranking Optimization for Human Alignment
abstract
Large language models (LLMs) often contain misleading content, emphasizing the need to align them with human values to ensure secure AI systems. Reinforcement learning from human feedback (RLHF) has been employed to achieve this alignment. However, it encompasses two main drawbacks: (1) RLHF exhibits complexity, instability, and sensitivity to hyperparameters in contrast to SFT. (2) Despite massive trial-and-error, multiple sampling is reduced to pair-wise contrast, thus lacking contrasts from a macro perspective. In this paper, we propose Preference Ranking Optimization (PRO) as an efficient SFT algorithm to directly fine-tune LLMs for human alignment. PRO extends the pair-wise contrast to accommodate preference rankings of any length. By iteratively contrasting candidates, PRO instructs the LLM to prioritize the best response while progressively ranking the rest responses. In this manner, PRO effectively transforms human alignment into aligning the probability ranking of n responses generated by LLM with the preference ranking of humans towards these responses. Experiments have shown that PRO outperforms baseline algorithms, achieving comparable results to ChatGPT and human responses through automatic-based, reward-based, GPT-4, and human evaluations.
Feifan Song 0001, Bowen Yu 0002, Haiyang Yu 0003, Fei Huang 0002, Yongbin Li 0001, Houfeng Wang
AAAI4
2024 Scaling Data Diversity for Fine-Tuning Language Models in Human Alignment
abstract
Alignment with human preference prevents large language models (LLMs) from generating misleading or toxic content while requiring high-cost human feedback. Assuming resources of human annotation are limited, there are two different ways of allocating considered: more diverse PROMPTS or more diverse RESPONSES to be labeled. Nonetheless, a straightforward comparison between their impact is absent. In this work, we first control the diversity of both sides according to the number of samples for fine-tuning, which can directly reflect their influence. We find that instead of numerous prompts, more responses but fewer prompts better trigger LLMs for human alignment. Additionally, the concept of diversity for prompts can be more complex than responses that are typically quantified by single digits. Consequently, a new formulation of prompt diversity is proposed, further implying a linear correlation with the final performance of LLMs after fine-tuning. We also leverage it on data augmentation and conduct experiments to show its effect on different algorithms.
Feifan Song 0001, Bowen Yu 0002, Hao Lang, Haiyang Yu 0003, Fei Huang 0002, Houfeng Wang, Yongbin Li 0001
LREC/COLING4
2024 Tree-Instruct: A Preliminary Study of the Intrinsic Relationship between Complexity and Alignment
abstract
Training large language models (LLMs) with open-domain instruction data has yielded remarkable success in aligning to end tasks and human preferences. Extensive research has highlighted the importance of the quality and diversity of instruction data. However, the impact of data complexity, as a crucial metric, remains relatively unexplored from three aspects: (1)where the sustainability of performance improvements with increasing complexity is uncertain; (2)whether the improvement brought by complexity merely comes from introducing more training tokens; and (3)where the potential benefits of incorporating instructions from easy to difficult are not yet fully understood. In this paper, we propose Tree-Instruct to systematically enhance the instruction complexity in a controllable manner. By adding a specified number of nodes to instructions’ semantic trees, this approach not only yields new instruction data from the modified tree but also allows us to control the difficulty level of modified instructions. Our preliminary experiments reveal the following insights: (1)Increasing complexity consistently leads to sustained performance improvements of LLMs. (2)Under the same token budget, a few complex instructions outperform diverse yet simple instructions. (3)Curriculum instruction tuning might not yield the anticipated results; focusing on increasing complexity appears to be the key.
Yingxiu Zhao, Bowen Yu 0002, Binyuan Hui, Haiyang Yu 0003, Fei Huang 0002, Nevin Lianwen Zhang, Yongbin Li 0001
LREC/COLING4
2024 Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
abstract
Minzheng Wang, Longze Chen, Fu Cheng, Shengyi Liao, Xinghua Zhang, Bingli Wu, Haiyang Yu, Nan Xu, Lei Zhang, Run Luo, Yunshui Li, Min Yang, Fei Huang, Yongbin Li. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Minzheng Wang 0001, Longze Chen, Fu Cheng, Shengyi Liao, Xinghua Zhang 0001, Bingli Wu, Haiyang Yu 0003, Nan Xu 0004, Lei Zhang 0201, Run Luo, Yunshui Li, Min Yang 0007, Fei Huang 0002, Yongbin Li 0001
EMNLP7
2024 Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch
abstract
In this paper, we unveil that Language Models (LMs) can acquire new capabilities by assimilating parameters from homologous models without retraining or GPUs. We first introduce DARE to set most delta parameters (i.e., the disparity between fine-tuned and pre-trained parameters) to zeros without affecting the abilities of Supervised Fine-Tuning (SFT) LMs, which randomly **D**rops delta parameters with a ratio $p$ **A**nd **RE**scales the remaining ones by $1 / (1 - p)$ to approximate the original embeddings. Then, we use DARE as a versatile plug-in to sparsify delta parameters of multiple SFT homologous models for mitigating parameter interference and merge them into a single model by parameter fusing. We experiment with encoder- and decoder-based LMs, showing that: (1) SFT delta parameter value ranges are typically small (within 0.002) with extreme redundancy, and DARE can effortlessly eliminate 90% or even 99% of them; (2) DARE can merge multiple task-specific LMs into one LM with diverse capabilities. Notably, this phenomenon is more pronounced in large-scale LMs, where the merged LM reveals the potential to surpass the performance of any source LM, providing a new discovery. We also utilize DARE to create a merged LM that ranks first among models with 7 billion parameters on the Open LLM Leaderboard.
Bowen Yu 0002, Haiyang Yu 0003, Fei Huang 0002, Yongbin Li 0001
ICML3
2024 Self-Retrieval: End-to-End Information Retrieval with One Large Language Model
abstract
The rise of large language models (LLMs) has significantly transformed both the construction and application of information retrieval (IR) systems. However, current interactions between IR systems and LLMs remain limited, with LLMs merely serving as part of components within IR systems, and IR systems being constructed independently of LLMs. This separated architecture restricts knowledge sharing and deep collaboration between them. In this paper, we introduce Self-Retrieval, a novel end-to-end LLM-driven information retrieval architecture. Self-Retrieval unifies all essential IR functions within a single LLM, leveraging the inherent capabilities of LLMs throughout the IR process. Specifically, Self-Retrieval internalizes the retrieval corpus through self-supervised learning, transforms the retrieval process into sequential passage generation, and performs relevance assessment for reranking. Experimental results demonstrate that Self-Retrieval not only outperforms existing retrieval approaches by a significant margin, but also substantially enhances the performance of LLM-driven downstream applications like retrieval-augmented generation.
Qiaoyu Tang, Jiawei Chen 0011, Bowen Yu 0002, Yaojie Lu 0001, Cheng Fu 0003, Haiyang Yu 0003, Fei Huang 0002, Ben He 0001, Xianpei Han, Le Sun 0001, Yongbin Li 0001
NeurIPS7
2023 Diversify Question Generation with Retrieval-Augmented Style Transfer
abstract
Given a textual passage and an answer, humans are able to ask questions with various expressions, but this ability is still challenging for most question generation (QG) systems.Existing solutions mainly focus on the internal knowledge within the given passage or the semantic word space for diverse content planning.These methods, however, have not considered the potential of external knowledge for expression diversity.To bridge this gap, we propose RAST, a framework for Retrieval-Augmented Style Transfer, where the objective is to utilize the style of diverse templates for question generation.For training RAST, we develop a novel Reinforcement Learning (RL) based approach that maximizes a weighted combination of diversity reward and consistency reward.Here, the consistency reward is computed by a Question-Answering (QA) model, whereas the diversity reward measures how much the final output mimics the retrieved template.Experimental results show that our method outperforms previous diversity-driven baselines on diversity while being comparable in terms of consistency scores.Our code is available at https://github.com/gouqi666/RAST.
Qi Gou, Zehua Xia, Bowen Yu 0002, Haiyang Yu 0003, Fei Huang 0002, Yongbin Li 0001, Cam-Tu Nguyen
EMNLP4
2023 API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
abstract
Recent research has demonstrated that Large Language Models (LLMs) can enhance their capabilities by utilizing external tools.However, three pivotal questions remain unanswered: (1) How effective are current LLMs in utilizing tools?(2) How can we enhance LLMs' ability to utilize tools?(3) What obstacles need to be overcome to leverage tools?To address these questions, we introduce API-Bank, a groundbreaking benchmark, specifically designed for tool-augmented LLMs.For the first question, we develop a runnable evaluation system consisting of 73 API tools.We annotate 314 tool-use dialogues with 753 API calls to assess the existing LLMs' capabilities in planning, retrieving, and calling APIs.For the second question, we construct a comprehensive training set containing 1,888 tool-use dialogues from 2,138 APIs spanning 1,000 distinct domains.Using this dataset, we train Lynx, a tool-augmented LLM initialized from Alpaca.Experimental results demonstrate that GPT-3.5 exhibits improved tool utilization compared to GPT-3, while GPT-4 excels in planning.However, there is still significant potential for further improvement.Moreover, Lynx surpasses Alpaca's tool utilization performance by more than 26 pts and approaches the effectiveness of GPT-3.5.Through error analysis, we highlight the key challenges for future research in this field to answer the third question 1 .
Yingxiu Zhao, Bowen Yu 0002, Feifan Song 0001, Hangyu Li 0003, Haiyang Yu 0003, Fei Huang 0002, Yongbin Li 0001
EMNLP6
2023 Causal Document-Grounded Dialogue Pre-training
abstract
The goal of document-grounded dialogue (DocGD) is to generate a response by anchoring the evidence in a supporting document in accordance with the dialogue context.This entails four causally interconnected variables.While task-specific pre-training has significantly enhanced performances on numerous downstream tasks, existing DocGD methods still rely on general pre-trained language models without a specifically tailored pre-training approach that explicitly captures the causal relationships.To address this, we present the first causallycomplete dataset construction strategy for developing million-scale DocGD pre-training corpora.Additionally, we propose a causallyperturbed pre-training strategy to better capture causality by introducing perturbations on the variables and optimizing the overall causal effect.Experiments conducted on three benchmark datasets demonstrate that our causal pretraining yields substantial and consistent improvements in fully-supervised, low-resource, few-shot, and zero-shot settings 1 .
Yingxiu Zhao, Bowen Yu 0002, Bowen Li 0002, Haiyang Yu 0003, Jinyang Li 0003, Fei Huang 0002, Yongbin Li 0001, Nevin Lianwen Zhang
EMNLP4
2023 Coarse-To-Fine Knowledge Selection for Document Grounded Dialogs
abstract
Multi-document grounded dialogue systems (DGDS) belong to a class of conversational agents that answer users’ requests by finding supporting knowledge from a collection of documents. Most previous studies aim to improve the knowledge retrieval model or propose more effective ways to incorporate external knowledge into a parametric generation model. These methods, however, focus on retrieving knowledge from mono-granularity language units (e.g. passages, sentences, or spans in documents), which is not enough to effectively and efficiently capture precise knowledge in long documents. This paper proposes Re3G, which aims to optimize both coarse-grained knowledge retrieval and fine-grained knowledge extraction in a unified framework. Specifically, the former efficiently finds relevant passages in a retrieval-and-reranking process, whereas the latter effectively extracts finer-grain spans within those passages to incorporate into a parametric answer generation model (BART, T5). Experiments on DialDoc Shared Task demonstrate the effectiveness of our method.
Yeqin Zhang, Haomin Fu, Cheng Fu 0003, Haiyang Yu 0003, Yongbin Li 0001, Cam-Tu Nguyen
ICASSP4
2022 Layout-Aware Information Extraction for Document-Grounded Dialogue: Dataset, Method and Demonstration
abstract
Building document-grounded dialogue systems have received growing interest as documents convey a wealth of human knowledge and commonly exist in enterprises. Wherein, how to comprehend and retrieve information from documents is a challenging research problem. Previous work ignores the visual property of documents and treats them as plain text, resulting in incomplete modality. In this paper, we propose a Layout-aware document-level Information Extraction dataset, LIE, to facilitate the study of extracting both structural and semantic knowledge from visually rich documents (VRDs), so as to generate accurate responses in dialogue systems. LIE contains 62k annotations of three extraction tasks from 4,061 pages in product and official documents, becoming the largest VRD-based information extraction dataset to the best of our knowledge. We also develop benchmark methods that extend the token-based language model to consider layout features like humans. Empirical results show that layout is critical for VRD-based extraction, and system demonstration also verifies that the extracted knowledge can help locate the answers that users care about.
Zhenyu Zhang 0006, Bowen Yu 0002, Haiyang Yu 0003, Tingwen Liu, Cheng Fu 0003, Chengguang Tang, Jian Sun 0021, Yongbin Li 0001
ACM Multimedia3
2020 Bridging Text and Knowledge with Multi-Prototype Embedding for Few-Shot Relational Triple Extraction
abstract
Current supervised relational triple extraction approaches require huge amounts of labeled data and thus suffer from poor performance in few-shot settings.However, people can grasp new knowledge by learning a few instances.To this end, we take the first step to study the few-shot relational triple extraction, which has not been well understood.Unlike previous single-task few-shot problems, relational triple extraction is more challenging as the entities and relations have implicit correlations.In this paper, We propose a novel multi-prototype embedding network model to jointly extract the composition of relational triples, namely, entity pairs and corresponding relations.To be specific, we design a hybrid prototypical learning mechanism that bridges text and knowledge concerning both entities and relations.Thus, implicit correlations between entities and relations are injected.Additionally, we propose a prototype-aware regularization to learn more representative prototypes.Experimental results demonstrate that the proposed method can improve the performance of the few-shot triple extraction.
Haiyang Yu 0003, Ningyu Zhang 0001, Shumin Deng, Hongbin Ye, Wei Zhang 0127, Huajun Chen
COLING1