Yanhua Yu

dblp:124/5763 · DBLP profile ↗
← Back
24ranked-venue papers
4as first author
23since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Hierarchical Reinforcement Learning with Augmented Step-Level Transitions for LLM Agents
abstract
Large language model (LLM) agents have demonstrated strong capabilities in complex interactive decision-making tasks.However, existing LLM agents typically rely on increasingly long interaction histories, resulting in high computational cost and limited scalability.In this paper, we propose STEP-HRL, a hierarchical reinforcement learning (HRL) framework that enables step-level learning by conditioning only on single-step transitions rather than full interaction histories.STEP-HRL structures tasks hierarchically, using completed subtasks to represent global progress of overall task.By introducing a local progress module, it also iteratively and selectively summarizes interaction history within each subtask to produce a compact summary of local progress.Together, these components yield augmented step-level transitions for both high-level and low-level policies.Experimental results on ScienceWorld and ALFWorld benchmarks consistently demonstrate that STEP-HRL substantially outperforms baselines in terms of performance and generalization while reducing token usage.
Shuai Zhen, Yanhua Yu, Ruopei Guo, Yang Deng 0002
ACL (1)2
2026 From Trajectories to States: State-Constrained Plan Execution for LLM Agents
Yanhua Yu, Nan Cheng 0001, Ruopei Guo, Ruowei Yin, Xidian Wang
ICIC (8)2
2026 Addressing graph heterogeneity and heterophily from a spectral perspective
Kangkang Lu 0002, Yanhua Yu, Ruopei Guo, Zhiyong Huang 0010, Yunshan Ma 0002, Meiyu Liang, Xiting Qin, Yimeng Ren 0001, Tat-Seng Chua
Neurocomputing2
2026 A Deep Reinforcement Learning-Based Fast Scheduling Approach for Dynamic Incremental Flows in Time Sensitive Networking
abstract
With the rapid advancement of intelligent and connected vehicles, in-vehicle networks face increasing demands for high-capacity data processing and transmission. Traditional static scheduling methods in Time Sensitive Network (TSN) can no longer satisfy the stringent real-time requirements of dynamic environments. The scheduling of time-sensitive traffic is typically addressed through offline computation, producing a mathematically constrained schedule that is subsequently deployed to end nodes or switches for configuration. However, such static deployment fails to accommodate dynamic scenarios. Consequently, this study proposes a dynamic incremental scheduling framework for TSN based on Deep Reinforcement Learning (DRL). The proposed framework integrates static global optimization with rapid dynamic adaptability. In the static phase, a Deep Q-Network (DQN) learns the priority ordering of time-sensitive flow scheduling under the given network state, thereby guiding a heuristic scheduler to efficiently allocate link resources. In the dynamic phase, an incremental scheduling strategy locally reconfigures resources for newly introduced flows without affecting existing schedules, thereby enhancing responsiveness and feasibility. Furthermore, this study proposes a load-balancing-guided routing set generation method that jointly optimizes path hop count and link load, enabling coordinated scheduling and routing. Comparative analysis under various network load scenarios demonstrates that the proposed method significantly reduces the response time of dynamic flow scheduling by 35% compared to the Satisfiability Modulo Theories (SMT) algorithm, while maintaining scheduling feasibility.
Feng Luo 0008, Yingpeng Tong, Yanhua Yu, Zhouping Zhang
IEEE Internet Things J.4
2025 LightPROF: A Lightweight Reasoning Framework for Large Language Model on Knowledge Graph
abstract
Large Language Models (LLMs) have impressive capabilities in text understanding and zero-shot reasoning. However, delays in knowledge updates may cause them to reason incorrectly or produce harmful results. Knowledge Graphs (KGs) provide rich and reliable contextual information for the reasoning process of LLMs by structurally organizing and connecting a wide range of entities and relations. Existing KG-based LLM reasoning methods only inject KGs' knowledge into prompts in a textual form, ignoring its structural information. Moreover, they mostly rely on close-source models or open-source models with large parameters, which poses challenges to high resource consumption. To address this, we propose a novel Lightweight and efficient Prompt learning-ReasOning Framework for KGQA (LightPROF), which leverages the full potential of LLMs to tackle complex reasoning tasks in a parameter-efficient manner. Specifically, LightPROF follows a “Retrieve-Embed-Reason” process, first accurately, and stably retrieving the corresponding reasoning graph from the KG through retrieval module. Next, through a Transformer-based Knowledge Adapter, it finely extracts and integrates factual and structural information from the KG, then maps this information to the LLM’s token embedding space, creating an LLM-friendly prompt to be used by the LLM for the final reasoning. Additionally, LightPROF only requires training Knowledge Adapter and can be compatible with any open-source LLM. Extensive experiments on two public KGQA benchmarks demonstrate that LightPROF achieves superior performance with small-scale LLMs. Furthermore, LightPROF shows significant advantages in terms of input token count and reasoning time.
Tu Ao, Yanhua Yu, Yang Deng 0002, Zirui Guo, Liang Pang 0001, Pinghui Wang, Tat-Seng Chua, Xiao Zhang 0046
AAAI2
2025 R2DQG: A Quality Meets Diversity Framework for Question Generation over Knowledge Bases
abstract
The task of Knowledge-Based Question Generation (KBQG) involves generating natural language questions from structured knowledge sources, posing unique challenges in balancing linguistic diversity and semantic relevance. Existing models often focus on maximizing surface-level similarity to ground-truth questions, neglecting the need for diverse syntactic forms and leading to semantic drift during generation. To overcome these challenges, we propose Refine-Reinforced Diverse Question Generation (R2DQG), a two-phase framework leveraging a generation-then-refinement paradigm. The Generator first constructs a diverse set of expressive templates using dependency parse tree similarity, capturing a wide range of syntactic patterns and styles. These templates guide the creation of question drafts, ensuring both diversity and semantic relevance. In the second phase, a Corrector module refines the drafts to mitigate semantic drift and enhance overall coherence and quality. Experiments on public datasets show that R2DQG outperforms state-of-the-art models in generating diverse, contextually accurate questions. Moreover, synthetic datasets generated by R2DQG enhance downstream QA performance, underscoring the practical utility of our approach.
Yimeng Ren 0001, Yanhua Yu, Lizi Liao, Yuhu Shang, Kangkang Lu 0002, Mingliang Yan
IJCAI2
2025 DeepMolTex: Deep Alignment of Molecular Graphs with Large Language Models via Mixture of Modality Experts
abstract
Recent advances in Molecular Graph-Language Models (MGLMs) have demonstrated promising capabilities in molecular understanding tasks. However, existing approaches face critical limitations: (1) shallow alignment methods which employ identical processing modules for both modalities, resulting in compromised expressiveness and catastrophic forgetting of pre-trained language capabilities; and (2) over-reliance on high-level molecular representations that inadequately capture fine-grained structural information essential for comprehensive molecular understanding. To address these challenges, we present DeepMolTex, a novel framework for Deep fusion of Molecular structure and Textual representations across multiple scales. Our approach introduces a Mixture of Modality Experts (MoME) architecture that facilitates deep alignment between molecular graph features and large language models while preserving language capabilities, and a multi-scale graph projector that extracts and aligns molecular features at atom, motif, and molecule levels. Experimental results demonstrate that DeepMolTex significantly outperforms existing methods on fundamental molecular understanding tasks, including molecule description generation and IUPAC name prediction, while effectively preserving the language capabilities of the pre-trained LLM.
Mingliang Yan, Yanhua Yu, Ruochi Zhang, Zhiyuan Liu 0010, Ruicheng Zhang, Yimeng Ren 0001, Kangkang Lu 0002, Zhiyong Huang 0010, Feng Luo 0004
ACM Multimedia2
2025 HiGraph-LLM: Hierarchical Graph Encoding and Integration with Large Language Models
Yanhua Yu, Xidian Wang, Kangkang Lu 0002, Tu Ao, Mingliang Yan, Liang Pang 0001, Pinghui Wang, Tat-Seng Chua
PRICAI2
2025 A centralized discovery-based method for integrating Data Distribution Service and Time-Sensitive Networking for In-Vehicle Networks
Feng Luo 0008, Yanhua Yu
Ad Hoc Networks3
2024 Improving Expressive Power of Spectral Graph Neural Networks with Eigenvalue Correction
abstract
In recent years, spectral graph neural networks, characterized by polynomial filters, have garnered increasing attention and have achieved remarkable performance in tasks such as node classification. These models typically assume that eigenvalues for the normalized Laplacian matrix are distinct from each other, thus expecting a polynomial filter to have a high fitting ability. However, this paper empirically observes that normalized Laplacian matrices frequently possess repeated eigenvalues. Moreover, we theoretically establish that the number of distinguishable eigenvalues plays a pivotal role in determining the expressive power of spectral graph neural networks. In light of this observation, we propose an eigenvalue correction strategy that can free polynomial filters from the constraints of repeated eigenvalue inputs. Concretely, the proposed eigenvalue correction strategy enhances the uniform distribution of eigenvalues, thus mitigating repeated eigenvalues, and improving the fitting capacity and expressive power of polynomial filters. Extensive experimental results on both synthetic and real-world datasets demonstrate the superiority of our method.
Kangkang Lu 0002, Yanhua Yu, Hao Fei 0001, Zixuan Yang 0001, Zirui Guo, Meiyu Liang, Mengran Yin, Tat-Seng Chua
AAAI2
2024 Hop-based Heterogeneous Graph Transformer
abstract
The Graph Transformer (GT) has shown significant ability in processing graph-structured data, addressing limitations in graph neural networks, such as over-smoothing and over-squashing. However, the implementation of GT in real-world heterogeneous graphs (HGs) with complex topology continues to present numerous challenges. Firstly, a challenge arises in designing a tokenizer that is compatible with heterogeneity. Secondly, the complexity of the transformer hampers the acquisition of high-order neighbor information in HGs. In this paper, we propose a novel Hop-based Heterogeneous Graph Transformer (H2Gormer) framework, paving a promising path for HGs to benefit from the capabilities of Transformers. We propose a Heterogeneous Hop-based Token Generation module to obtain high-order information in a flexible way. Specifically, to enrich the fine-grained heterogeneous semantics of each token, we propose a tailored multi-relational encoder to encode the hop-based neighbors. In this way, the resulting token embeddings are input to the Hop-based Transformer to obtain node representations, which are then combined with position embeddings to obtain the final encoding. Extensive experiments on four datasets are conducted to demonstrate the effectiveness of H2Gormer.
Zixuan Yang 0001, Xiao Wang 0017, Yanhua Yu, Kangkang Lu 0002, Zirui Guo, Xiting Qin, Yunshan Ma 0002, Tat-Seng Chua
ECAI3
2024 Adversary and Attention Guided Knowledge Graph Reasoning Based on Reinforcement Learning
Yanhua Yu, Xiuxiu Cai, Ang Ma, Yimeng Ren 0001, Shuai Zhen, Kangkang Lu 0002, Zhiyong Huang 0010, Tat-Seng Chua
KSEM (5)1
2024 EE-LCE: An Event Extraction Framework Based on LLM-Generated CoT Explanation
Yanhua Yu, Yunshan Ma 0002, Kangkang Lu 0002, Zhiyong Huang 0010, Tat-Seng Chua
KSEM (1)1
2024 Information-Controllable Graph Contrastive Learning for Recommendation
abstract
In the evolving landscape of recommender systems, Graph Contrastive Learning (GCL) has become a prominent method for enhancing recommendation performance by alleviating the issue of data sparsity. However, existing GCL-based recommendations often overlook the control of shared information between the contrastive views. In this paper, we initially analyze and experimentally demonstrate these methods often lead to the issue of augmented representation collapse, where the representations between views become excessively similar, diminishing their distinctiveness. To address this issue, we propose the Information-Controllable Graph Contrastive Learning (IGCL) framework, a novel approach that focuses on optimizing the shared information between views to include as much relevant information for the recommendation task as possible while maintaining an appropriate level. In particular, we design the Collaborative Signals Enhanced Augmentation module to infuse the augmented representation with rich, task-relevant collaborative signals. Furthermore, the Information-Controllable Contrastive Learning module is designed to direct control over the magnitude of shared information between the contrastive views to avoid over-similarity. Extensive experiments on three public datasets demonstrate the effectiveness of IGCL, showcasing significant improvements in performance and the capability to alleviate augmented representation collapse.
Zirui Guo, Yanhua Yu, Kangkang Lu 0002, Zixuan Yang 0001, Liang Pang 0001, Tat-Seng Chua
RecSys2
2024 Can Small Language Models be Good Reasoners for Sequential Recommendation?
abstract
Large language models (LLMs) open up new horizons for sequential recommendations, owing to their remarkable language comprehension and generation capabilities. However, there are still numerous challenges that should be addressed to successfully implement sequential recommendations empowered by LLMs. Firstly, user behavior patterns are often complex, and relying solely on one-step reasoning from LLMs may lead to incorrect or task-irrelevant responses. Secondly, the prohibitively resource requirements of LLM (e.g., ChatGPT-175B) are overwhelmingly high and impractical for real sequential recommender systems. In this paper, we propose a novel Step-by-step knowLedge dIstillation fraMework for recommendation (SLIM), paving a promising path for sequential recommenders to enjoy the exceptional reasoning capabilities of LLMs in a "slim" (i.e. resource-efficient) manner. We introduce CoT prompting based on user behavior sequences for the larger teacher model. The rationales generated by the teacher model are then utilized as labels to distill the downstream smaller student model (e.g., LLaMA2-7B). In this way, the student model acquires the step-by-step reasoning capabilities in recommendation tasks. We encode the generated rationales from the student model into a dense vector, which empowers recommendation in both ID-based and ID-agnostic scenarios. Extensive experiments demonstrate the effectiveness of SLIM over state-of-the-art baselines, and further analysis showcasing its ability to generate meaningful recommendation reasoning at affordable costs.
Changxin Tian, Binbin Hu, Yanhua Yu, Zhiqiang Zhang 0012, Jun Zhou 0011, Liang Pang 0001, Xiao Wang 0017
WWW4
2024 Cross-view hypergraph contrastive learning for attribute-aware recommendation
Ang Ma, Yanhua Yu, Chuan Shi 0001, Zirui Guo, Tat-Seng Chua
Inf. Process. Manag.2
2023 Neural Reconstruction through Scattering Media with Forward and Backward Losses
abstract
Reconstructing an object behind a scattering medium needs to tackle different light scatterings inside the medium and in the free space. Major approaches, e.g., diffuse optical tomography and non-line-of-sight imaging, address either of the light scatterings. Confocal diffuse tomography (CDT) considers the media acting as a diffuse kernel on the free-space scattering and recovers the object by deconvolving the inside-medium scattering. Inspired by CDT, we present a Neural De-Scatterer to solve this challenging problem. We exploit neural implicit fields to represent the the free-space scattering and use a multilayer perceptron (MLP) to learn the density and albedo of the hidden object. Furthermore, we tailor a bi-directional training strategy to optimize the MLP with forward and backward losses and employ hash encoding for memory and computation efficiency. The Neural De-Scatterer enables us to reconstruct objects at arbitrary resolution. Comprehensive experiments with synthetic and real measurements demonstrate that our Neural De-Scatterer outperforms state-of-the-art methods. Our data and code are publicly available.
Yuehan Wang, Suan Xia, Ruiqian Li, Xingyue Peng, Yanhua Yu, Jingyi Yu 0001
ICCP6
2023 Enhancing Non-line-of-sight Imaging via Learnable Inverse Kernel and Attention Mechanisms
abstract
Recovering information from non-line-of-sight (NLOS) imaging is a computationally-intensive inverse problem. Most physics-based NLOS imaging methods address the complexity of this problem by assuming three-bounce reflections and no self-occlusion. However, these assumptions may break down for objects with large depth variations, preventing physics-based algorithms from accurately reconstructing the details and high-frequency information. On the other hand, while learning-based methods can avoid these assumptions, they may struggle to reconstruct details without specific designs due to the spectral bias of neural networks. To overcome these issues, we propose a novel approach that enhances physics-based NLOS imaging methods by introducing a learnable inverse kernel in the Fourier domain and using an attention mechanism to improve the neural network to learn high-frequency information. Our method is evaluated on publicly available and new synthetic datasets, demonstrating its commendable performance compared to prior physics-based and learning-based methods, especially for objects with large depth variations. Moreover, our approach generalizes well to real data and can be applied to tasks such as classification and depth reconstruction. We will make our code and dataset publicly available: https://sci2020.github.io.
Yanhua Yu, Zi Wang 0019, Binbin Huang 0004, Yuehan Wang, Xingyue Peng, Suan Xia, Ruiqian Li
ICCV1
2023 Deep Unsupervised Momentum Contrastive Hashing for Cross-modal Retrieval
abstract
Unsupervised cross-modal hashing (UCMH) methods often start from the similarity of sample features and design a reconstruction loss to achieve similarity preservation. However, these methods suffer from inaccurate similarity problems, be-cause different feature representations may share similar semantic information. In this paper, we propose Deep Unsupervised Momentum Contrastive Hashing (DUMCH). Specifically, we introduce momentum contrastive learning for unsupervised cross-modal hashing, which allows us to flexibly define a robust loss by comparing positive and negative samples. Moreover, in order to achieve similarity retention of hash codes in Hamming space and fully utilize the potential of contrastive learning in Hamming space, we remove the L2 normalization corresponding to cosine similarity and design a novel normalization method called hash normalization, which has been proved to greatly improve the model performance. We conducted extensive experiments on three datasets, and the experimental results demonstrate the superiority of DUMCH.
Kangkang Lu 0002, Yanhua Yu, Meiyu Liang, Xiaowen Cao 0003, Zehua Zhao, Mengran Yin, Zhe Xue
ICME2
2023 Intent-aware Recommendation via Disentangled Graph Contrastive Learning
abstract
Graph neural network (GNN) based recommender systems have become one of the mainstream trends due to the powerful learning ability from user behavior data. Understanding the user intents from behavior data is the key to recommender systems, which poses two basic requirements for GNN-based recommender systems. One is how to learn complex and diverse intents especially when the user behavior is usually inadequate in reality. The other is different behaviors have different intent distributions, so how to establish their relations for a more explainable recommender system. In this paper, we present the Intent-aware Recommendation via Disentangled Graph Contrastive Learning (IDCL), which simultaneously learns interpretable intents and behavior distributions over those intents. Specifically, we first model the user behavior data as a user-item-concept graph, and design a GNN based behavior disentangling module to learn the different intents. Then we propose the intent-wise contrastive learning to enhance the intent disentangling and meanwhile infer the behavior distributions. Finally, the coding rate reduction regularization is introduced to make the behaviors of different intents orthogonal. Extensive experiments demonstrate the effectiveness of IDCL in terms of substantial improvement and the interpretability.
Xiao Wang 0017, Xiangzhou Huang, Yanhua Yu, Haoyang Li 0001, Mengdi Zhang 0002, Zirui Guo, Wei Wu 0014
IJCAI4
2022 HiddenPose: Non-Line-of-Sight 3D Human Pose Estimation
abstract
Nearly all existing human pose estimation techniques address the problem under the line-of-sight (LOS) setting. Many real-life applications such as rescue missions and autonomous driving, in contrast, require estimating the pose of hidden subjects. In this paper, we present a non-line-of-sight (NLOS) pose estimator, which produces a skeletal representation of hidden human poses. A brute-force approach would first conduct albedo reconstruction of a hidden subject and then apply LOS pose estimation. We show that such an implementation does not effectively exploit features unique to NLOS and subsequently yields artifacts such as missing joints. We instead first generate a comprehensive NLOS human pose dataset of 19 subjects under 9 motions. We then present a spatially aware deep learning technique based on convolutional neural networks that explicitly employ NLOS features. Comprehensive experiments on both synthetic and real data show that our new estimator is both effective and robust and can be seamlessly integrated into learning-based NLOS scene reconstruction. Our HiddenPose transient dataset contains synthetic transients with ground-truths of the volumes and the joints and real-world transients captured from our NLOS imaging system. Extensive assessments demonstrate that the HiddenPose transient dataset is valuable for effective NLOS research. We will make our data and code publicly available.
Yanhua Yu, Zhengqing Pan, Xingyue Peng, Ruiqian Li, Yuehan Wang, Jingyi Yu 0001
ICCP2
2022 Ensemble Multi-Relational Graph Neural Networks
abstract
It is well established that graph neural networks (GNNs) can be interpreted and designed from the perspective of optimization objective. With this clear optimization objective, the deduced GNNs architecture has sound theoretical foundation, which is able to flexibly remedy the weakness of GNNs. However, this optimization objective is only proved for GNNs with single-relational graph. Can we infer a new type of GNNs for multi-relational graphs by extending this optimization objective, so as to simultaneously solve the issues in previous multi-relational GNNs, e.g., over-parameterization? In this paper, we propose a novel ensemble multi-relational GNNs by designing an ensemble multi-relational (EMR) optimization objective. This EMR optimization objective is able to derive an iterative updating rule, which can be formalized as an ensemble message passing (EnMP) layer with multi-relations. We further analyze the nice properties of EnMP layer, e.g., the relationship with multi-relational personalized PageRank. Finally, a new multi-relational GNNs which well alleviate the over-smoothing and over-parameterization issues are proposed. Extensive experiments conducted on four benchmark datasets well demonstrate the effectiveness of the proposed model.
Yanhua Yu, Mengdi Zhang 0002, Yuji Yang, Wei Wu 0014
IJCAI3
2022 Automatic Graph Generation for Document-Level Relation Extraction
abstract
Relation extraction, one of important natural language processing tasks, aims to find out the semantic relations between entities in text. It has been raised to the document level recently, which is a complex task that requires a logical inference to extract relations from multiple entities in intra-and inter-sentence. Existing methods mainly utilize co-references or syntactic trees to establish static document-level graphs. Different from these heuristic models which may not always yield the optimal structures, we propose a novel model that can automatically construct task-specific graphs. We regard the process of establishing a graph as a sequence problem and utilize a gated recurrent unit (GRU) recurrent neural networks (RNNs) to construct a document graph. Afterwards, we utilize the graph convolutional networks (GCNs) to calculate the node representation. Particularly, our model exhibits comparable performance to state-of-the-art models on the large-scale human annotated document-level relation extraction dataset (DocRED).
Yanhua Yu, Fangting Shen, Shengli Yang, Ang Ma
IJCNN1
2020 Diversified Multiple Instance Learning for Document-Level Multi-Aspect Sentiment Classification
abstract
Neural Document-level Multi-aspect Sentiment Classification (DMSC) usually requires a lot of manual aspect-level sentiment annotations, which is time-consuming and laborious.As document-level sentiment labeled data are widely available from online service, it is valuable to perform DMSC with such free document-level annotations.To this end, we propose a novel Diversified Multiple Instance Learning Network (D-MILN), which is able to achieve aspect-level sentiment classification with only document-level weak supervision.Specifically, we connect aspect-level and document-level sentiment by formulating this problem as multiple instance learning, providing a way to learn aspect-level classifier from the back propagation of document-level supervision.Two diversified regularizations are further introduced in order to avoid the overfitting on document-level signals during training.Diversified textual regularization encourages the classifier to select aspect-relevant snippets, and diversified sentimental regularization prevents the aspect-level sentiments from being overly consistent with document-level sentiment.Experimental results on TripAdvisor and BeerAdvocate datasets show that D-MILN remarkably outperforms recent weaklysupervised baselines, and is also comparable to the supervised method.
Yunjie Ji, Hao Liu 0026, Bolei He, Xinyan Xiao, Hua Wu 0003, Yanhua Yu
EMNLP (1)6