Guoyi Li

dblp:65/8028 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2026 Knowledge-Enhanced Image Captioning with Adaptive Graph-based Multimodal Alignment and LLM
abstract
Image captioning is crucial for multimodal understanding, bridging visual content and natural language. Despite recent advancements in Large Multimodal Models (LMMs), when faced with unseen entities or scenes in the open world, even when attempting to leverage learned knowledge, models still struggle with vague and inaccurate descriptions, and may even generate knowledge hallucinations. A key reason is that the model fails to effectively integrate knowledge with visual information, limiting its understanding of visual content. Thus, we propose Adaptive Knowledge Graph-guided Multimodal Alignment (AKGMA) for image captioning, which enhances semantic understanding in open-world scenes through visual knowledge reasoning, reducing knowledge hallucinations and improving caption quality. It consist three key components: Entity-guided Knowledge Aligner (EKA), Adaptive Knowledge Graph Construction (AKGC), and Scene-Context Knowledge Adapter (SCKA). EKA connects visual entities to knowledge graphs, providing structured knowledge to a small language model, which interacts with a visual encoder to acquire visual knowledge. AKGC uses reinforcement learning to build image-relevant subgraphs to optimize knowledge prompts and improve knowledge hallucinations. SCKA leverages scene graph annotations to extract visual contextual knowledge and inject it into Large Language Models (LLMs), ensuring the generated descriptions are consistent with the image's details. Additionally, we introduce UniKnowCap, a new image knowledge description dataset spanning various open-world knowledge domains, designed to evaluate the knowledge accuracy and detail consistency of model-generated descriptions. Extensive experiments show our model outperforms baselines across multiple metrics.
Guoyi Li, Die Hu 0004, Zhongjiang Yao, Wei Mi, Zongzhen Liu, Xiaodan Zhang 0004, Honglei Lyu
AAAI1
2025 Semantic Reshuffling with LLM and Heterogeneous Graph Auto-Encoder for Enhanced Rumor Detection
abstract
Social media is crucial for information spread, necessitating effective rumor detection to curb misinformation’s societal effects. Current methods struggle against complex propagation influenced by bots, coordinated accounts, and echo chambers, which fragment information and increase risks of misjudgments and model vulnerability. To counteract these issues, we introduce a new rumor detection framework, the Narrative-Integrated Metapath Graph Auto-Encoder (NIMGA). This model consists of two core components: (1) Metapath-based Heterogeneous Graph Reconstruction. (2) Narrative Reordering and Perspective Fusion. The first component dynamically reconstructs propagation structures to capture complex interactions and hidden pathways within social networks, enhancing accuracy and robustness. The second implements a dual-agent mechanism for viewpoint distillation and comment narrative reordering, using LLMs to refine diverse perspectives and semantic evolution, revealing patterns of information propagation and latent semantic correlations among comments. Extensive testing confirms our model outperforms existing methods, demonstrating its effectiveness and robustness in enhancing rumor representation through graph reconstruction and narrative reordering.
Guoyi Li, Die Hu 0004, Zongzhen Liu, Xiaodan Zhang 0004, Honglei Lyu
COLING1
2025 WebSurfer: Enhancing LLM Agents with Web-Wise Feedback for Web Navigation
abstract
As the Internet’s complexity and information volume surge, the need for efficient web automation becomes critical. Traditional web agents struggle with redundant web content, which disrupts their understanding of the environment. They also face inefficiencies in multi-task scenarios due to handcrafted exemplars and encounter error accumulation in long-horizon tasks, exacerbated by web-specific complexities like nested structures and interactive elements. To address these issues, we introduce WebSurfer, a novel web agent designed to filter, learn, and adapt in complex environments. WebSurfer refines task-oriented states for clearer observations and employs an exemplar retrieval and ordering strategy to enhance LLMs’ understanding and adaptability to current tasks. Notably,WebSurfer features a novel web-wise insight feedback mechanism that enables continuous adaptation and strategy refinement. Evaluations demonstrate that WebSurfer outperforms state-of-the-art (SOTA) methods on realistic tasks, achieving higher accuracy and enhancing longterm adaptability.
Die Hu 0004, Jingguo Ge, Weitao Tang, Guoyi Li, Liangxiong Li, Bingzhen Wu
ICASSP4
2025 Emotion-aware Structural Enhancement Graph Auto-Encoder for Rumor Detection
abstract
Social media is a key channel for information dissemination, making effective rumor detection essential to mitigate misinformation’s societal impact. Although large language models excel in inference and text generation, they struggle with understanding propagation relationships and complex reasoning tasks like rumor detection. Existing methods mainly rely on textual information and event propagation structures, but provocative comments and unreliable interactions increase propagation uncertainty. To address these challenges, we propose an Emotionally-Aware Structural Enhancement Graph Auto-Encoder (EASE-GARD) to improve rumor representations. Our method begins by enhancing the textual representation of responses through the generation of emotive adversarial comments. It then generates and differentiates false local propagation relationships (fabricated forwards and reciprocations) to reduce propagation uncertainties. A graph auto-encoder captures contextual features and global structural information and recalculates forwarding probabilities among responses. Extensive experiments show our model performs best on all three datasets, excelling in effectiveness and robustness.
Guoyi Li, Zhongjiang Yao, Die Hu 0004, Yingrui Xu, Xiaodan Zhang 0004, Honglei Lyu
ICASSP1
2025 Entity Graph Alignment and Visual Reasoning for Multimodal Fake News Detection
abstract
The rise of multimodal fake news threatens reliable information dissemination by exploiting multiple modalities to create deceptive, engaging content, significantly impacting society safety. Existing methods still face challenges in cross-modal alignment (e.g., semantic inconsistencies, complex visual-semantic relations) and are vulnerable to low-quality or noisy samples. To address these, we propose Cross-Modal Alignment with Visual Reasoning Prompting (CMA-VRP) for multimodal fake news detection. Specifically, we model text and image entities with graphs to capture fine-grained semantic interactions and enhance cross-modal consistency through graph contrastive learning. Unlike methods relying on shallow image features (e.g., edges, textures), we leverage large language models (LLMs) and large vision-language models (LVLMs) to capture deep visual-semantic attributes related to reasoning (e.g., actions, scenes). Based on graph modeling and visual reasoning features, we perform graph-based cross-modal semantic fusion to unify textual and visual representations and cross-modal cycle alignment to align modality distributions by reducing semantic discrepancies, filtering modality-specific noise, and extracting invariant representations across domains. These steps enable the model to obtain semantically consistent and modality-invariant features. Extensive experiments demonstrate that our model outperforms existing methods in multimodal fake news detection and shows strong robustness against noisy samples.
Guoyi Li, Die Hu 0004, Xiaomeng Fu, Qirui Tang, Yulei Wu, Xiaodan Zhang 0004, Honglei Lyu
ACM Multimedia1
2025 Zero-Shot Multimodal Fact-Checking with Conceptual Reasoning
abstract
In multimodal fact-checking, advanced large multimodal models (LMMs) struggle to capture and integrate the complex relationships between text and images. A potential solution is to generate reasoning support text to optimize reasoning and integrate evidence. However, existing generation approaches rely heavily on high-quality data annotations for training, which are costly and limited in scalability, hindering responsiveness to evolving misinformation. To address these issues, we propose CoReS, a novel zero-shot multimodal fact-checking model based on Conceptual Reasoning Support-leveraging key concepts from evidence to guide the reasoning process and improve decision-making. This model includes a reasoning support text generation module that extracts key concepts (critical elements that significantly impact the judgment outcome) from raw textual evidence via retrieval and filtering. By using a Conceptual Reasoning LM, CoReS generates reasoning support texts framed around core key concepts that are semantically consistent with multimodal evidence, linking key clues, thus replacing redundant and complex evidence for fact-checking. The reasoning support texts generated by CoReS effectively distill complex evidence relationships and integrate important reasoning information, allowing the judgment model to provide clear and accurate judgments. Evaluations on benchmark datasets and the new multi-domain MultiVerify dataset demonstrate that CoReS excels in accuracy, generalization, and scalability.
Guoyi Li, Die Hu 0004, Qirui Tang, Xiaomeng Fu, Yulei Wu, Xiaodan Zhang 0004, Honglei Lyu
ACM Multimedia1
2024 CRDA: Content Risk Drift Assessment of Large Language Models through Adversarial Multi-Agent Interaction
abstract
As Large Language Models (LLMs) continue to enhance their capabilities in multi-agent collaborative applications, the unpredictability of the generative content risks has intensified. Particularly in ongoing interaction scenarios with users, it remains unclear whether there is generative content risk drift over time. In this context, "drift risk" refers to the trend of progressively intensified content risk that emerges during sustained adversarial interactions among LLM agents. Additionally, the high cost associated with constructing complex adversarial environments for agents impedes the transferability of current assessment methods for LLMs to multi-agent adversarial scenarios. In this paper, we introduce a low-cost and lightweight framework for assessing content risk drift of LLMs, named CRDA. This framework, bypassing the need for constructing complex adversarial environments, offers a method that integrates roles and responses memory to guide automatically multi-round adversarial interactions among LLM agents, that is, multiple agents as avatars of a single LLM. In this approach, LLM agents enable the analysis of content risk drift of this LLM. Moreover, we explore the impact of restricted roles and the unsafe content with negative viewpoints in responses memory on the content risk drift of LLMs. Considering the rapid advancement of Chinese LLM capabilities, this study selects real adversarial topics in Chinese and assesses content risk drift of five representative Chinese LLMs. The research finds that these LLMs exhibit significant content risk drift even after a certain safety alignment, showing an initial increase followed by a gradual decrease. As the adversarial process progresses, under restricted roles, agents more effectively breach the model's safety alignment, leading to content risk drift of the LLM. The content drift risk assessment can be quantified specifically by measuring the deterioration rate at which LLM agents deteriorate from positive to negative and analyzing the underlying trends during the automatically multi-round adversarial interactions. In restricted and general roles adversarial interactions, all agents of five Chinese LLMs exhibit an overall average increase of 31.5% and 16.38% in the cumulative deterioration rate respectively by the 10th round, compared to the baseline no-roles adversarial interactions. Finally, we hope that the framework and findings presented in this paper will offer valuable insights for research on safety alignment in LLM agents during adversarial processes.
Zongzhen Liu, Guoyi Li, Bingkang Shi, Xiaodan Zhang 0004, Jingguo Ge, Yulei Wu, Honglei Lyu
IJCNN2
2024 Finite-Time Multi-Parameter Smoothing Distributed Algorithm for Delayed Multi-Agent Systems Subject to Switching Topology
abstract
This brief tackles the finite-time consensus and distributed optimization issues in multi-agent systems (MASs) affected by time-varying delays (TDs). It centers on scenarios where the agents' communication network topology is characterized by a switching undirected graph. Distributed time-varying optimization involves a collective effort by multiple agents to collaboratively minimize the aggregate of local objective functions that change over time. This process is subject to time-dependent equality constraints and relies solely on information available at the local level and from neighboring agents. First, a finite-time multi-parameter smooth distributed algorithm with input TDs is proposed to solve the optimization problem of MASs with arbitrary switching topologies. Secondly, combined with the Lyapunov stability theory, the states of each agent can achieve consensus and asymptotically track the optimal solution trajectory within a finite-time is proved. At last, an illustrative simulation example of target tracking with MASs is given to verify the effectiveness and applicability of the theoretical results.
Jiahao Leng, Qishui Zhong, Sheng Han 0002, Guoyi Li, Gangming Zhang
SMC4
2024 Multimodal Fake News Detection Based on Chain-of-Thought Prompting Large Language Models
abstract
The rapid rise of social networks has led to a proliferation of fake news, especially those with images. The combination of images and text may confuse users and cause even more negative impact. Exisiting methods for fake news detection either require expert knowledge or large amounts of labeled data. In addition, these methods fails to clarify which part of the multimodal information is misleading or why. In this paper, we present a simple yet efficient Chain-of-thought Prompting method for Multimodal Fake News Detection (CP-FEND). It first finds the closest demonstration samples of the news posts to be detected by a KNN-based approach. Afterwards, we design a logical prompt method including Examination, Inference and Determination stages to guide Large Language Model (LLM) to automatically construct reasoning processes for the authenticity of the samples. Finally, LLM are prompted to derive the authenticity of multimodal news with the guidance of samples and Chain-of-Thought reasoning. A reflective verification is performed to further improve the detection performance through comprehensive evaluation of the original responses. Extentsive experiments on two public datasets have demonstrated the superiority of our method over existing methods.
Yingrui Xu, Jingguo Ge, Guangxu Lyu, Guoyi Li, Hui Li 0098
SMC4
2022 Viewport-Oriented Panoramic Image Inpainting
abstract
Panoramic images are usually viewed through Head Mounted Displays (HMDs), which renders only a narrow field of view from the raw panoramic image. This distinctive viewing feature has largely been ignored when inpainting panoramic images. To address this issue, we propose a viewport-oriented generative adversarial panoramic image inpainting network in this paper. For capturing the distorted features accurately in the generating process of equirectangular projection (ERP) panoramic image, a latitude-adaptive feature fusion module is devised to aggregate the latitude-level features in ERP image and less-distorted patch-level viewport-domain features. Furthermore, a novel cross-domain discriminator is proposed to force the inpainting network to generate more plausible results in viewports. Extensive experiments show that our model achieves better performance compared to the baseline methods, especially in the viewport images.
Zhuoyi Shang, Yanwei Liu 0001, Guoyi Li, Jingbo Miao, Jinxia Liu, Liming Wang 0001
ICIP3
2022 A Heterogeneous Propagation Graph Model for Rumor Detection Under the Relationship Among Multiple Propagation Subtrees
Guoyi Li, Yulei Wu, Xiaodan Zhang 0004, Wei Zhou 0019, Honglei Lyu
ECML/PKDD (2)1
2020 Real-time anomaly detection framework using a support vector regression for the safety monitoring of commercial aircraft
Hyunseong Lee, Guoyi Li, Ashwin Rai, Aditi Chattopadhyay
Adv. Eng. Informatics2
2001 ABR services for overcoming misbehaving sources in a heterogeneous environment
Guoyi Li, Ian Li-Jin Thng, Lawrence Wai-Choong Wong
J. Netw. Comput. Appl.1