Yupu Hao

dblp:362/7775 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 44% Learning paradigms · 17% Information extraction and text analysis · 11%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
chain-of-thought reasoning
1.012026
Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do · ACL (1) 2026
Machine learning › Learning paradigms
continual learning
1.012026
Spectral Disentanglement: Rank-Aware Task Adaptation for Rehearsal-free Continual Learning in LLMs · ACL (1) 2026
Natural language and speech › Language models and text generation › chain-of-thought reasoning
multimodal chain-of-thought
1.012026
Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do · ACL (1) 2026
Computer vision › Vision and language
multimodal reasoning
1.012026
Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do · ACL (1) 2026
Machine learning › Learning paradigms › continual learning
rehearsal-free continual learning
1.012026
Spectral Disentanglement: Rank-Aware Task Adaptation for Rehearsal-free Continual Learning in LLMs · ACL (1) 2026
Natural language and speech › Language models and text generation › large language model evaluation
LLM-as-a-judge
0.912025
Evaluating Personalized Tool-Augmented LLMs from the Perspectives of Personalization and Proactivity · ACL (1) 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.912025
CITI: Enhancing Tool Utilizing Ability in Large Language Models Without Sacrificing General Performance · AAAI 2025
Natural language and speech › Language models and text generation › LLM agents
tool learning
0.912025
CITI: Enhancing Tool Utilizing Ability in Large Language Models Without Sacrificing General Performance · AAAI 2025
Natural language and speech › Language models and text generation › LLM agents
tool use
0.912025
CITI: Enhancing Tool Utilizing Ability in Large Language Models Without Sacrificing General Performance · AAAI 2025
Natural language and speech › Information extraction and text analysis › event extraction
event detection
0.712023
Event Ontology Completion with Hierarchical Structure Evolution Networks · EMNLP 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning › ontology
event ontology
0.712023
Event Ontology Completion with Hierarchical Structure Evolution Networks · EMNLP 2023
Natural language and speech › Information extraction and text analysis › event extraction
event type induction
0.712023
Event Ontology Completion with Hierarchical Structure Evolution Networks · EMNLP 2023
Natural language and speech › Language models and text generation
large language model
0.312026
Spectral Disentanglement: Rank-Aware Task Adaptation for Rehearsal-free Continual Learning in LLMs · ACL (1) 2026
Natural language and speech › Language models and text generation › large language model › large language model adaptation
personalization
0.312025
Evaluating Personalized Tool-Augmented LLMs from the Perspectives of Personalization and Proactivity · ACL (1) 2025

Methods — techniques the papers use, named apart from their topics

spectral disentanglement · 1.0rank-aware task adaptation · 1.0chain-of-thought prompting · 1.0key-point-based evaluation · 0.9gradient-based importance score · 0.9fine-tuning · 0.9Mixture-of-LoRA · 0.9in-context learning · 0.7hierarchical linking · 0.7contrastive clustering · 0.7
YearPublicationVenuePosition
2026 Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do
abstract
Zhuoran Jin, Kejian Zhu, Hongbang Yuan, Yupu Hao, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhuoran Jin, Kejian Zhu, Hongbang Yuan, Yupu Hao, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001
ACL (1)4
2026 Spectral Disentanglement: Rank-Aware Task Adaptation for Rehearsal-free Continual Learning in LLMs
abstract
Huanxuan Liao, Shizhu He, Yupu Hao, Yequan Wang, Wenhao Teng, Xiangwen Liao, Jun Zhao, Kang Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Huanxuan Liao, Shizhu He, Yupu Hao, Yequan Wang, Wenhao Teng, Xiangwen Liao, Jun Zhao 0001, Kang Liu 0001
ACL (1)3
2025 CITI: Enhancing Tool Utilizing Ability in Large Language Models Without Sacrificing General Performance
abstract
Tool learning enables Large Language Models (LLMs) to interact with the external environment by invoking tools, enriching the accuracy and capability scope of LLMs. However, previous works predominantly focus on improving the model's tool-utilizing accuracy and the ability to generalize to new, unseen tools, excessively forcing LLMs to adjust specific tool-invoking pattern without considering the harm to the model's general performance. This deviates from the actual applications and original intention of integrating tools to enhance the model. To tackle this problem, we dissect the capability trade-offs by examining the hidden representation changes and the gradient-based importance score of the model's components. Based on the analysis result, we propose a Component Importance-based Tool-utilizing ability Injection method (CITI). According to the gradient-based importance score of different components, it alleviates the capability conflicts caused by the fine-tuning process by applying distinct training strategies to different components. CITI applies Mixture-Of-LoRA (MOLoRA) for important components. Meanwhile, it fine-tunes the parameters of a few components deemed less important in the backbone of the LLM, while keeping other parameters frozen. CITI can effectively enhance the model's tool-utilizing capability without excessively compromising its general performance. Experimental results demonstrate that our approach achieves outstanding performance across a range of evaluation metrics.
Yupu Hao, Zhuoran Jin, Huanxuan Liao, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001
AAAI1
2025 Evaluating Personalized Tool-Augmented LLMs from the Perspectives of Personalization and Proactivity
abstract
Personalized tool utilization is essential for aligning large language models (LLMs) with user preference in interaction scenarios with various tools. However, most of the current benchmarks primarily focus on either personalization of text generation or direct tool-utilizing, without considering both. In this work, we introduce a novel benchmark ETAPP for evaluating personalized tool invocation, establishing a sandbox environment, and a comprehensive dataset of 800 testing cases covering diverse user profiles. To improve the accuracy of our evaluation, we propose a key-point-based LLM evaluation method, mitigating biases in the LLM-as-a-judge system by manually annotating key points for each test case and providing them to LLM as the reference. Additionally, we evaluate the excellent LLMs and provide an in-depth analysis. Furthermore, we investigate the impact of different tool-invoking strategies on LLMs’ personalization performance and the effects of fine-tuning in our task. The effectiveness of our preference-setting and key-point-based evaluation method is also validated. Our findings offer insights into improving personalized LLM agents. Our code is available at https://github.com/hypasd-art/ETAPP.
Yupu Hao, Zhuoran Jin, Huanxuan Liao, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001
ACL (1)1
2025 SKIntern: Internalizing Symbolic Knowledge for Distilling Better CoT Capabilities into Small Language Models
abstract
Small Language Models (SLMs) are attracting attention due to the high computational demands and privacy concerns of Large Language Models (LLMs). Some studies fine-tune SLMs using Chains of Thought (CoT) data distilled from LLMs, aiming to enhance their reasoning ability. Furthermore, Some CoT distillation methods introduce external symbolic knowledge into the generation process to improve the limited knowledge memory, reasoning ability and out-of-domain (OOD) generalization of SLMs. However, the introduction of symbolic knowledge increases computational overhead and introduces potential noise. In this paper, we introduce SKIntern, an innovative approach that empowers SLMs to internalize symbolic knowledge and few-shot examples gradually through a progressive fine-tuning process, guided by a predefined linear decay schedule under curriculum learning. By efficiently internalizing knowledge, SKIntern reduces computational overhead and speeds up the reasoning process by focusing solely on the question during inference. It outperforms state-of-the-art baselines by over 5%, while reducing inference costs (measured in FLOPs) by up to 4\times across a wide range of SLMs in both in-domain (ID) and out-of-domain (OOD) tasks. Our code will be available at https://github.com/Xnhyacinth/SKIntern.
Huanxuan Liao, Shizhu He, Yupu Hao, Yuanzhe Zhang, Jun Zhao 0001, Kang Liu 0001
COLING3
2023 Event Ontology Completion with Hierarchical Structure Evolution Networks
abstract
Traditional event detection methods require predefined event schemas.However, manually defining event schemas is expensive and the coverage of schemas is limited.To this end, some works study the event type induction (ETI) task, which discovers new event types via clustering.However, the setting of ETI suffers from two limitations: event types are not linked into the existing hierarchy and have no semantic names.In this paper, we propose a new research task named Event Ontology Completion (EOC), which aims to simultaneously achieve event clustering, hierarchy expansion and type naming.Furthermore, we develop a HierarchicAL STructure EvOlution Network (HALTON) for this new task.Specifically, we first devise a Neighborhood Contrastive Clustering module to cluster unlabeled event instances.Then, we propose a Hierarchy-Aware Linking module to incorporate the hierarchical information for event expansion.Finally, we generate meaningful names for new types via an In-Context Learning-based Naming module.Extensive experiments indicate that our method achieves the best performance, outperforming the baselines by 8.23%, 8.79% and 8.10% of ARI score on three datasets 1 .
Yupu Hao, Yubo Chen 0001, Kang Liu 0001, Jiexin Xu, Huaijun Li, Xiaojian Jiang, Jun Zhao 0001
EMNLP2