Zheyong Xie

dblp:339/1743 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2026
0009-0009-7453-5781ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Optimizing Generative Ranking Relevance via Reinforcement Learning in Xiaohongshu Search
abstract
Ranking relevance is a fundamental task in search engines, aiming to identify the items most relevant to a given user query. Traditional relevance models typically produce scalar scores or directly predict relevance labels, limiting both interpretability and the modeling of complex relevance signals. Inspired by recent advances in Chain-of-Thought (CoT) reasoning for complex tasks, we investigate whether explicit reasoning can enhance both interpretability and performance in relevance modeling. However, existing reasoning-based Generative Relevance Models (GRMs) primarily rely on supervised fine-tuning on large amounts of human-annotated or synthetic CoT data, which often leads to limited generalization. Moreover, domain-agnostic, free-form reasoning tends to be overly generic and insufficiently grounded, limiting its potential to handle the diverse and ambiguous cases prevalent in open-domain search. In this work, we formulate relevance modeling in Xiaohongshu search as a reasoning task and introduce a Reinforcement Learning (RL)-based training framework to enhance the grounded reasoning capabilities of GRMs. Specifically, we incorporate practical business-specific relevance criteria into the multi-step reasoning prompt design and propose Stepwise Advantage Masking (SAM), a lightweight process-supervision strategy which facilitates effective learning of these criteria through improved credit assignment. To enable industrial deployment, we further distill the large-scale RL-tuned model to a lightweight version suitable for real-world search systems. Extensive offline evaluations and online A/B tests demonstrate that our approach consistently delivers significant improvements across key relevance and business metrics, validating its effectiveness, robustness, and practicality for large-scale industrial search systems.
Ziyang Zeng, Heming Jing, Jindong Chen, Yige Sun, Zheyong Xie, Shaosheng Cao, Yao Hu 0002
KDD (1)9
2025 Pet-Bench: Benchmarking the Abilities of Large Language Models as E-Pets in Social Network Services
abstract
As interest in using Large Language Models for interactive and emotionally rich experiences grows, virtual pet companionship emerges as a novel yet underexplored application. Existing approaches focus on basic pet role-playing interactions without systematically benchmarking LLMs for comprehensive companionship. In this paper, we introduce PET-BENCH, a dedicated benchmark that evaluates LLMs across both self-interaction and human-interaction dimensions. Unlike prior work, PET-BENCH emphasizes self-evolution and developmental behaviors alongside interactive engagement, offering a more realistic reflection of pet companionship. It features diverse tasks such as intelligent scheduling, memory-based dialogues, and psychological conversations, with over 7,500 interaction instances designed to simulate pet behaviors. Evaluation of 28 LLMs reveals significant performance variations linked to model size and inherent capabilities, underscoring the need for specialized optimization in this domain. PET-BENCH serves as a foundational resource for benchmarking pet-related LLM abilities and advancing emotionally immersive human-pet interactions.
Hongcheng Guo, Zheyong Xie, Shaosheng Cao, Boyang Wang 0006, Weiting Liu 0001, Zheyu Ye, Zhoujun Li 0001, Zuozhu Liu, Wei Lu 0011
CIKM2
2025 PaRT: Enhancing Proactive Social Chatbots with Personalized Real-Time Retrieval
abstract
Social chatbots have become essential companions in daily scenarios ranging from emotional support to personal interaction. However, conventional chatbots with passive response mechanisms usually rely on users to initiate or sustain dialogues by bringing up new topics, resulting in diminished engagement and shortened dialogue duration. In this paper, we present PaRT, a novel framework enabling context-aware proactive dialogues for social chatbots through personalized real-time retrieval and generation. Specifically, PaRT first integrates user profiles and dialogue context into a large language model (LLM), which is initially prompted to refine user queries and recognize underlying intents for the upcoming conversation. Guided by refined intents, the LLM generates personalized dialogue topics as targeted queries to retrieve relevant passages from RedNote. Finally, we prompt LLMs with summarized passages to generate knowledge-grounded and engagement-optimized responses. Our approach has been running stably in a real-world production environment for more than 30 days, achieving a 21.77% improvement in the average duration of dialogues.
Zihan Niu, Zheyong Xie, Shaosheng Cao, Chonggang Lu, Zheyu Ye, Tong Xu 0001, Zuozhu Liu, Yan Gao 0017, Jia Chen 0003, Yao Hu 0002
SIGIR2
2024 Retrieve-Plan-Generation: An Iterative Planning and Answering Framework for Knowledge-Intensive LLM Generation
abstract
Despite the significant progress of large language models (LLMs) in various tasks, they often produce factual errors due to their limited internal knowledge.Retrieval-Augmented Generation (RAG), which enhances LLMs with external knowledge sources, offers a promising solution.However, these methods can be misled by irrelevant paragraphs in retrieved documents.Due to the inherent uncertainty in LLM generation, inputting the entire document may introduce off-topic information, causing the model to deviate from the central topic and affecting the relevance of the generated content.To address these issues, we propose the Retrieve-Plan-Generation (RPG) framework.RPG generates plan tokens to guide subsequent generation in the plan stage.In the answer stage, the model selects relevant fine-grained paragraphs based on the plan and uses them for further answer generation.This plan-answer process is repeated iteratively until completion, enhancing generation relevance by focusing on specific topics.To implement this framework efficiently, we utilize a simple but effective multi-task prompt-tuning method, enabling the existing LLMs to handle both planning and answering.We comprehensively compare RPG with baselines across 5 knowledge-intensive generation tasks, demonstrating the effectiveness of our approach.1
Yuanjie Lyu, Zihan Niu, Zheyong Xie, Chao Zhang 0096, Tong Xu 0001, Yang Wang 0001, Enhong Chen
EMNLP3
2024 Knowledge-Enhanced Multi-perspective Incongruity Perception Network for Multimodal Sarcasm Detection
abstract
Recent years have witnessed the urgent request for multi-modal sarcasm detection in social media platforms. Though large efforts have been made with significant progress, prior arts may fail to fully integrate the commonsense knowledge, and struggle to model the incongruity within implicit meanings of multimodal cues. To address these limitations, we propose a novel Knowledge-Enhanced Multi-perspective Incongruity Perception network, named KEMIP. Specifically, we adopt generative language models to produce captions and commonsense for image and text respectively for comprehensive understanding. Subsequently, to exploit the essential cues from multiple perspectives, we customize an incongruity perception module, which utilizes several fusion networks to capture both literal and implicit inconsistencies. Afterwards, an ensemble weighting gate is employed to integrate the result of individual perspective. Experiments on a public multimodal sarcasm detection benchmark demonstrate the superiority of our proposed KEMIP framework.
Zihan Niu, Zheyong Xie, Tong Xu 0001, Xiangfeng Wang 0005, Yao Hu 0002, Enhong Chen
ICME2
2024 Speak From Heart: An Emotion-Guided LLM-Based Multimodal Method for Emotional Dialogue Generation
abstract
Recent advancements in Large Language Models~(LLMs) have greatly enhanced the generation capabilities of dialogue systems. However, progress on emotional expression during dialogues might be still limited, especially when capturing and processing the multimodal cues for emotional expression. Therefore, it is urgent to fully adapt the multimodal understanding ability and transferability of LLMs to enhance the emotional-oriented multimodal processing capabilities. To that end, in this paper, we propose a novel Emotion-Guided Multimodal Dialogue model based on LLM, termed ELMD. Specifically, to enhance the emotional expression ability of LLMs, our ELMD customizes an emotional retrieval module, which mainly provides appropriate response demonstration for LLM in understanding emotional context. Subsequently, a two-stage training strategy is proposed, founded on previous demonstration support, to support uncovering nuanced emotions behind multimodal information and constructing natural responses. Comprehensive experiments demonstrate the effectiveness and superiority of ELMD.
Chenxiao Liu, Zheyong Xie, Sirui Zhao, Tong Xu 0001, Minglei Li 0001, Enhong Chen
ICMR2
2023 Comprehending the Gossips: Meme Explanation in Time-Sync Video Comment via Multimodal Cues
abstract
Recent years have witnessed the booming of online social media platforms with embracing the popular service called “Time-Sync Comment”, which supports the viewers to share their time-sync opinions along with video content. In this way, we observe that numerous semantically-altered terms, or “Memes”, were created by niche users to express their unique ideas and emotions, and further attracted a large group of viewers with better activity and enthusiasm. Unfortunately, since the memes were created based on domain-specific knowledge and semantically varied depending on the multimodal context in videos, newcomers may fail to comprehend the semantic connotation of memes, which may severely impair their user-experiences. To deal with this issue, in this article, we propose a novel meme explanation framework, called ProMDE, to automatically capture and comprehend the memes in time-sync comments, which could further benefit the viewers with meme explanation service. Specifically, we first iteratively reconstruct the original time-sync comments compared with visual embedding to detect the semantically-altered terms as meme candidates. Afterward, based on the guides from the domain-specific corpus, visual and textual features will be fused to represent the context-aware multimodal cues. Moreover, to accurately describe the commonly-seen homophones in memes, i.e., they have the same pronunciation but different word-spelling expressions, we integrate the phonetic symbols as an additional modality to enhance the framework. Finally, we utilize a Transformer-based decoder to generate the natural language explanation for captured memes. Extensive experiments on a large real-world dataset prove that our framework could significantly outperform several state-of-the-art baseline methods, demonstrating the efficacy of modeling multimodal context and pronunciation for meme detection and explanation.
Zheyong Xie, Weidong He, Tong Xu 0001, Chen Zhu 0003, Ping Yang 0010, Enhong Chen
ACM Trans. Asian Low Resour. Lang. Inf. Process.1