Yuhan Liu 0023

dblp:125/8141-23 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0003-2912-750XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains
abstract
Vision-language models (VLMs) achieve remarkable success in single-image tasks.However, real-world scenarios often involve intricate multi-image inputs, leading to a notable performance decline as models struggle to disentangle critical information scattered across complex visual features.In this work, we propose Focus-Centric Visual Chain, a novel paradigm that enhances VLMs' perception, comprehension, and reasoning abilities in multi-image scenarios.To facilitate this paradigm, we propose Focus-Centric Data Synthesis, a scalable bottom-up approach for synthesizing high-quality data with elaborate reasoning paths.Through this approach, We construct VISC-150K, a large-scale dataset with reasoning data in the form of Focus-Centric Visual Chain, specifically designed for multi-image tasks.Experimental results on seven multi-image benchmarks demonstrate that our method achieves average performance gains of 3.16% and 2.24% across two distinct model architectures, without compromising the general vision-language capabilities.Our study represents a significant step toward more robust and capable vision-language systems that can handle complex visual scenarios: VISC. * Corresponding authors.Which of the following images contains the same object as the first image and shares the same attribute weight?
Juntian Zhang, Chuanqi Cheng, Yuhan Liu 0023, Wei Liu 0302, Jian Luan 0001, Rui Yan 0001
ACL (1)3
2025 More is not always better? Enhancing Many-Shot In-Context Learning with Differentiated and Reweighting Objectives
abstract
Xiaoqing Zhang, Ang Lv, Yuhan Liu, Flood Sung, Wei Liu, Jian Luan, Shuo Shang, Xiuying Chen, Rui Yan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xiaoqing Zhang 0017, Ang Lv, Yuhan Liu 0023, Flood Sung, Wei Liu 0302, Jian Luan 0001, Shuo Shang, Xiuying Chen, Rui Yan 0001
ACL (1)3
2025 Beyond Static Testbeds: An Interaction-Centric Agent Simulation Platform for Dynamic Recommender Systems
abstract
Evaluating and iterating upon recommender systems is crucial, yet traditional A/B testing is resource-intensive, and offline methods struggle with dynamic user-platform interactions.While agent-based simulation is promising, existing platforms often lack a mechanism for user actions to dynamically reshape the environment.To bridge this gap, we introduce RecInter, a novel agent-based simulation platform for recommender systems featuring a robust interaction mechanism.In RecInter platform, simulated user actions (e.g., likes, reviews, purchases) dynamically update item attributes in real-time, and introduced Merchant Agents can reply, fostering a more realistic and evolving ecosystem.High-fidelity simulation is ensured through Multidimensional User Profiling module, Advanced Agent Architecture, and LLM fine-tuned on Chain-of-Thought (CoT) enriched interaction data.Our platform achieves significantly improved simulation credibility and successfully replicates emergent phenomena like Brand Loyalty and the Matthew Effect.Experiments demonstrate that this interaction mechanism is pivotal for simulating realistic system evolution, establishing our platform as a credible testbed for recommender systems research: RecInter.
Juntian Zhang, Yuhan Liu 0023, Guojun Yin, Rui Yan 0001
EMNLP3
2025 The Truth Becomes Clearer Through Debate! Multi-Agent Systems with Large Language Models Unmask Fake News
abstract
In today's digital environment, the rapid propagation of fake news via social networks poses significant social challenges. Most existing detection methods either employ traditional classification models, which suffer from low interpretability and limited generalization capabilities, or craft specific prompts for large language models (LLMs) to produce explanations and results directly, failing to leverage LLMs' reasoning abilities fully. Inspired by the saying that ''truth becomes clearer through debate,'' our study introduces a novel multi-agent system with LLMs named TruEDebate (TED) to enhance the interpretability and effectiveness of fake news detection. TED employs a rigorous debate process inspired by formal debate settings. Central to our approach are two innovative components: the DebateFlow Agents and the InsightFlow Agents. The DebateFlow Agents organize agents into two teams, where one supports and the other challenges the truth of the news. These agents engage in opening statements, cross-examination, rebuttal, and closing statements, simulating a rigorous debate process akin to human discourse analysis, allowing for a thorough evaluation of news content. Concurrently, the InsightFlow Agents consist of two specialized sub-agents: the Synthesis Agent and the Analysis Agent. The Synthesis Agent summarizes the debates and provides an overarching viewpoint, ensuring a coherent and comprehensive evaluation. The Analysis Agent, which includes a role-aware encoder and a debate graph, integrates role embeddings and models the interactions between debate roles and arguments using an attention mechanism, providing the final judgment.Our extensive experiments on two datasets, ARG-EN and ARG-CN, demonstrate that the TED framework surpasses traditional methods across various metrics and, more importantly, enhances interpretable fake news detection by illuminating logical reasoning and structured debate processes leading to accurate conclusions.We release our code to support Information systems that use structured debate within responsible information systems for improved decision-making.
Yuhan Liu 0023, Yuxuan Liu 0009, Xiaoqing Zhang 0017, Xiuying Chen, Rui Yan 0001
SIGIR1
2025 SAGraph: A Large-Scale Social Graph Dataset with Comprehensive Context for Influencer Selection in Marketing
abstract
Influencer marketing campaign success heavily depends on identifying key opinion leaders who can effectively leverage their credibility and reach to promote products or services.The selection of influencers is vital for boosting brand visibility, fostering consumer trust, and driving sales.While traditional research often simplifies complex factors like user attitudes, interaction frequency, and advertising content, into simple numerical values.However, this reductionist approach fails to capture the dynamic nature of influencer marketing effectiveness.To bridge this gap, we present SAGraph, a novel comprehensive dataset from Weibo that captures multi-dimensional marketing campaign data across six product domains.The dataset encompasses 345,039 user profiles with their complete interaction histories, including 1.3M comments and 554K reposts across 44K posts, providing unprecedented granularity in influencer marketing dynamics.SAGraph uniquely integrates user profiles, content features, and temporal interaction patterns, enabling in-depth analysis of influencer marketing mechanisms.Experimental results using both traditional baselines and state-of-the-art large language models (LLMs) demonstrate the crucial role of content analysis in predicting advertising effectiveness.Our findings reveal that LLM-based approaches achieve superior performance in understanding and predicting campaign success, opening new avenues for data-driven influencer marketing strategies.We hope that this dataset will inspire further research: https
Xiaoqing Zhang 0017, Yuhan Liu 0023, Zhenxing Hu, Xiuying Chen, Rui Yan 0001
SIGIR2
2024 IAD: In-Context Learning Ability Decoupler of Large Language Models in Meta-Training
abstract
Large Language Models (LLMs) exhibit remarkable In-Context Learning (ICL) ability, where the model learns tasks from prompts consisting of input-output examples. However, the pre-training objectives of LLMs often misalign with ICL objectives. They’re mainly pre-trained with methods like masked language modeling and next-sentence prediction. On the other hand, ICL leverages example pairs to guide the model in generating task-aware responses such as text classification and question-answering tasks. The basic pre-training task-related capabilities can sometimes overshadow or conflict with task-specific subtleties required in ICL. To address this, we propose an In-context learning Ability Decoupler (IAD). The model aims to separate the ICL ability from the general ability of LLMs in the meta-training phase, where the ICL-related parameters are separately tuned to adapt for ICL tasks. Concretely, we first identify the parameters that are suitable for ICL by transference-driven gradient importance. We then propose a new max-margin loss to emphasize the separation of the general and ICL abilities. The loss is defined as the difference between the output of ICL and the original LLM, aiming to prevent the overconfidence of the LLM. By meta-training these ICL-related parameters with max-margin loss, we enable the model to learn and adapt to new tasks with limited data effectively. Experimental results show that IAD’s capability yields state-of-the-art performance on benchmark datasets by utilizing only 30% of the model’s parameters. Ablation study and detailed analysis prove the separation of the two abilities.
Yuhan Liu 0023, Xiuying Chen, Ji Zhang 0011, Rui Yan 0001
LREC/COLING1
2024 From Skepticism to Acceptance: Simulating the Attitude Dynamics Toward Fake News
Yuhan Liu 0023, Xiuying Chen, Xiaoqing Zhang 0017, Ji Zhang 0011, Rui Yan 0001
IJCAI1
2024 Bridging the Space Gap: Unifying Geometry Knowledge Graph Embedding with Optimal Transport
abstract
Knowledge Graph Embedding (KGE) is a critical field aiming to transform the elements of knowledge graphs (KGs) into continuous spaces, offering great potential for structured data representation. In contemporary KGE research, the utilization of either hyperbolic or Euclidean space for knowledge graph Embedding is a common practice. However, knowledge graphs encompass diverse geometric data structures, including chains and hierarchies, whose hybrid nature exceeds the capacity of a single embedding space to capture effectively. This paper introduces a novel and highly effective approach called Unified Geometry Knowledge Graph Embedding (UniGE) to address the challenge of representing diverse geometric data in KGs. UniGE stands out as a novel KGE method that seamlessly integrates KGE in both Euclidean and hyperbolic geometric spaces. We introduce an embedding alignment method and fusion strategy, which harnesses optimal transport techniques and the Wasserstein barycenter method. Furthermore, we offer a comprehensive theoretical analysis to substantiate the superiority of our approach, as evident from a more robust error bound. To substantiate the strength of UniGE, we conducted comprehensive experiments on three benchmark datasets. The results consistently demonstrate that UniGE outperforms state-of-the-art methods, aligning with the conclusions drawn from our theoretical analysis.
Yuhan Liu 0023, Zelin Cao, Ji Zhang 0011, Rui Yan 0001
WWW1
2018 Clustering based on grid and local density with priority-based expansion for multi-density data
Shaoqun Dong, Yuhan Liu 0023, Lianbo Zeng, Chaoshui Xu, Tingying Zhou
Inf. Sci.3