Yingji Li

dblp:292/2914 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0002-3575-1395ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Breaking the Passive Learning Trap: An Active Perception Strategy for Human Motion Prediction
abstract
Forecasting 3D human motion is an important embodiment of fine-grained understanding and cognition of human behavior by artificial agents. Current approaches excessively rely on implicit network modeling of spatiotemporal relationships and motion characteristics, falling into the passive learning trap that results in redundant and monotonous 3D coordinate information acquisition while lacking actively guided explicit learning mechanisms. To overcome these issues, we propose an Active Perceptual Strategy (APS) for human motion prediction, leveraging quotient space representations to explicitly encode motion properties while introducing auxiliary learning objectives to strengthen spatio-temporal modeling. Specifically, we first design a data perception module that projects poses into the quotient space, decoupling motion geometry from coordinate redundancy. By jointly encoding tangent vectors and Grassmann projections, this module simultaneously achieves geometric dimension reduction, semantic decoupling, and dynamic constraint enforcement for effective motion pose characterization. Furthermore, we introduce a network perception module that actively learns spatio-temporal dependencies through restorative learning. This module deliberately masks specific joints or injects noise to construct auxiliary supervision signals. A dedicated auxiliary learning network is designed to actively adapt and learn from perturbed information. Notably, APS is model agnostic and can be integrated with different prediction models to enhance active perceptual.The experimental results demonstrate that our method achieves the new state-of-the-art, outperforming existing methods by large margins: 16.3% on H3.6M, 13.9% on CMU Mocap, and 10.1% on 3DPW.
Juncheng Hu 0002, Zijian Zhang 0009, Yingji Li, Kedi Lyu
AAAI5
2026 Enhancing Multimodal Large Language Models for Ancient Chinese Character Evolution Analysis via Glyph-Driven Fine-Tuning
abstract
In recent years, rapid advances in Multimodal Large Language Models (MLLMs) have increasingly stimulated research on ancient Chinese scripts.As the evolution of written characters constitutes a fundamental pathway for understanding cultural transformation and historical continuity, how MLLMs can be systematically leveraged to support and advance text evolution analysis remains an open and largely underexplored problem.To bridge this gap, we construct a comprehensive benchmark comprising 11 tasks and over 130,000 instances, specifically designed to evaluate the capability of MLLMs in analyzing the evolution of ancient Chinese scripts.We conduct extensive evaluations across multiple widely used MLLMs and observe that, while existing models demonstrate a limited ability in glyph-level comparison, their performance on core tasks-such as character recognition and evolutionary reasoning-remains substantially constrained.Motivated by these findings, we propose a glyph-driven fine-tuning framework (GEVO) that explicitly encourages models to capture evolutionary consistency in glyph transformations and enhances their understanding of text evolution.Experimental results show that even models at the 2B scale achieve consistent and comprehensive performance improvements across all evaluated tasks.To facilitate future research, we publicly release both the benchmark and the trained models 1 .
Rui Song 0008, Lida Shi, Ruihua Qi, Yingji Li, Hao Xu 0012
ACL (1)4
2026 Toward Robust In-Context Learning: Leveraging Out-of-distribution Proxies for Target Inaccessible Demonstration Retrieval
abstract
Although studies have demonstrated that Large Language Models (LLMs) can perform well on Out-of-Distribution (OOD) tasks, their advantage tends to diminish as the distribution shift becomes more severe. Consequently, researchers aim to retrieve distributionally similar and informative demonstrations from the available source domain to boost the inference capabilities of LLMs. However, in practical scenarios where the target domain is inaccessible, evaluating the unknown distribution is challenging, which indirectly impacts the quality of the selected demonstrations. To address this problem, we propose DOPA, a demonstration search framework that incorporates an OOD proxy to approximate the inaccessible target domain and guide the retrieval process. Building on proxy-based evaluation, DOPA further introduces a Mahalanobis distance-based global diversity constraint to ensure sufficient diversity among the retrieved demonstrations. Experimental results on multiple LLMs and tasks demonstrate that DOPA effectively enhances robustness in OOD settings.
Hao Xu 0012, Rite Bo, Fausto Giunchiglia, Yingji Li, Rui Song 0008
ACL (1)4
2026 Retrieval augmentation for out-of-distribution robustness in non-knowledge intensive in-context learning
Rui Song 0008, Yingji Li, Fausto Giunchiglia, Hao Xu 0012
Sci. China Inf. Sci.2
2026 Enhancing fairness in decision-making of natural language understanding systems: An intersectional bias debiasing model via information theory-based disentanglement
Yingji Li, Juncheng Hu 0002, Rui Song 0008, Liang Hu 0001
Expert Syst. Appl.1
2026 How GenAI tools influence the purchase intention of green products through the mediating role of emotional connection: Evidence from China
abstract
In the field of green product consumption, consumers tend to seek detailed information to assess product efficacy. In recent years, the advent of Generative Artificial Intelligence (GenAI) tools has significantly streamlined consumers’ access to relevant information on green products. However, as emotional factors are decisive in purchase decision-making, existing studies that predominantly focus on rational decision-making frequently overlook this crucial emotional dimension. Addressing this gap, this study adopts the extended emotion heuristic theory to examine the impact of GenAI tools on green product purchase preferences. Using the partial least squares structural equation model, green product consumption data were collected from multiple regions including Chongqing, Guangdong, Hunan, Hubei, Shanghai, Beijing, and others, between January and March 2025. A total of 717 valid responses were analysed using SPSS 28, Amos 28, and Smart PLS 4.0. The results reveal that certain characteristics of GenAI-generated content—specifically, quality (content relevance, content accuracy), communication style (personalisation, anthropomorphism), and serendipity—positively influence purchase intention for green products. Furthermore, emotional connection plays a partial mediating role. These findings extend the application of emotion heuristic theory in the context of artificial intelligence and highlight the significant role of emotional factors in fostering consumption intentions via GenAI tools. The results offer insights for green product marketers and GenAI tool developers to enhance content quality, communication methods, and additional functions, while also informing regulatory policymaking related to GenAI tools.
Xingpeng Zheng, Yue Xia, Yingji Li
Inf. Process. Manag.4
2026 Causal contrastive learning for generalizable graph classification
Baohang Wei, Yingji Li, Ying Wang 0009
Knowl. Based Syst.2
2025 Generalizable Graph Prompt Learning Framework with Model-level Prompt Injection and Two-Stage Prompt Tuning
abstract
Graph prompt learning represents a novel paradigm aimed at enhancing the performance of graph learning models on a variety of downstream tasks by providing specific graph prompts. Despite its promise, current graph prompt learning methods are limited by the following limitations. On the one hand, existing methods often rely on manually selected graph information or simple learnable vectors, which can introduce human biases and lack expressiveness. These methods also fall short in guiding models to induce historical prior knowledge and improve generalization. Furthermore, the direct end-to-end tuning strategy of prompts lacks a necessary gentle transition, which impacts model stability and generalization. To overcome these limitations, we introduce the generalizable graph prompt learning framework (GGPL), which incorporates model-level prompt injection and a two-stage prompt tuning strategy. GGPL focuses on encoding subgraph structures and attributes during pre-training and uses SimGRACE to predict subgraph similarities, enhancing the base model's generalization. The model-level prompt injection module, with its prompt embedding backbone and self-prompt generation, seamlessly integrates invariant knowledge. Our two-stage tuning strategy, including transition and task-specific tuning, ensures better guidance and stability. By designing learnable prompt tokens and fine-tuning them with task-specific information, GGPL enables the model to generalize more robustly to downstream tasks. We conduct extensive experiments on six benchmark datasets to verify the model's effectiveness.
Mingchen Sun, Jiahui Hou, Yingji Li, Ying Wang 0009
KDD (2)4
2025 BATED: Learning fair representation for Pre-trained Language Models via biased teacher-guided disentanglement
Yingji Li, Mengnan Du, Rui Song 0008, Mu Liu, Ying Wang 0009
Artif. Intell.1
2025 Causal keyword driven reliable text classification with large language model feedback
Rui Song 0008, Yingji Li, Mingjie Tian, Fausto Giunchiglia, Hao Xu 0012
Inf. Process. Manag.2
2025 Counterfactual contrastive learning for robust text classification based on word group search
Rui Song 0008, Fausto Giunchiglia, Yingji Li, Lida Shi, Hao Xu 0012
Inf. Sci.3
2025 KALD: A Knowledge Augmented multi-contrastive learning model for low resource abusive Language Detection
Rui Song 0008, Fausto Giunchiglia, Yingji Li, Jian Li 0080, Hao Xu 0012
Knowl. Based Syst.3
2024 TACIT: A Target-Agnostic Feature Disentanglement Framework for Cross-Domain Text Classification
abstract
Cross-domain text classification aims to transfer models from label-rich source domains to label-poor target domains, giving it a wide range of practical applications. Many approaches promote cross-domain generalization by capturing domaininvariant features. However, these methods rely on unlabeled samples provided by the target domains, which renders the model ineffective when the target domain is agnostic. Furthermore, the models are easily disturbed by shortcut learning in the source domain, which also hinders the improvement of domain generalization ability. To solve the aforementioned issues, this paper proposes TACIT, a target domain agnostic feature disentanglement framework which adaptively decouples robust and unrobust features by Variational Auto-Encoders. Additionally, to encourage the separation of unrobust features from robust features, we design a feature distillation task that compels unrobust features to approximate the output of the teacher. The teacher model is trained with a few easy samples that are easy to carry potential unknown shortcuts. Experimental results verify that our framework achieves comparable results to state-of-the-art baselines while utilizing only source domain data.
Rui Song 0008, Fausto Giunchiglia, Yingji Li, Mingjie Tian, Hao Xu 0012
AAAI3
2024 Mitigating social biases of pre-trained language models via contrastive self-debiasing with double data augmentation
Yingji Li, Mengnan Du, Rui Song 0008, Xin Wang 0035, Mingchen Sun, Ying Wang 0009
Artif. Intell.1
2024 Exploring Multiple Hypergraphs for Heterogeneous Graph Neural Networks
abstract
Graph neural networks have demonstrated significant power in learning graph representations for homogeneous networks. However, real-world network data can often be denoted by heterogeneous networks with different types of nodes and edges, such as social, traffic, and molecular networks. Network heterogeneity presents significant challenges for network analysis and mining. Motif-based hypergraphs preserve high-order proximity and capture composite semantic interactions. Because not all nodes and edges in the original network always exist in a specific hypergraph, it is essential that multiple motif-based hypergraphs are considered to enhance the network representation. Therefore, we propose a novel framework for exploring M ultiple M o tif-based H ypergraphs for H eterogeneous G raph N eural N etworks to learn network representations, named MoH-HGNN, which leverages hypergraph convolution and attention operations to capture complex connectivity patterns. Specifically, we conducted two levels of attention networks with hierarchical structures, namely hyperedge-level attention to learn the importance among different types of nodes and comprehensive semantic-level attention to capture the importance of different types of motif structures. We extensively experimented on four real-world datasets to verify the effectiveness of our proposed framework.
Ying Wang 0009, Yingji Li, Xin Wang 0035
Expert Syst. Appl.2
2024 Towards Domain-Aware Stable Meta Learning for Out-of-Distribution Generalization
abstract
Deep learning models are often trained on datasets that are limited in size and distribution, which may not fully represent the entire range of data encountered in practice. Thus, making deep learning models generalize to out-of-distribution data has received a significant amount of attention in recent studies due to the critical importance of this ability in real-world applications. Meta learning as an effective knowledge transfer paradigm, which learns a base model with high generalization ability to adapt to new data distributions by minimizing domain shifts across tasks during meta-training. However, most existing meta learning methods assume that the base model can access the labels of different domains, and this assumption is demanding in many real application scenarios. In addition, these methods focus on narrowing data-level domain shifts, while ignoring task-level domain shifts, which may lead to inadequate or even negative transfer. Inspired by human learners who use induction to learn and master new tasks, we propose a novel domain-aware meta learning framework for out-of-distribution generalization, termed SMLG. This framework enables the base model to generalize effectively to unseen domains without relying on domain-specific labels. Specifically, we develop a domain-aware transformation module to obtain meta representation and pseudo domain labels. As a result, the base model can be trained robustly without the need for direct domain label input. Furthermore, to investigate the impact of domain shifts at different levels, we introduce a joint loss function that combines cross-entropy with a domain alignment constraint. Extensive experiments on benchmark datasets demonstrate the efficacy of our framework.
Mingchen Sun, Yingji Li, Ying Wang 0009, Xin Wang 0035
ACM Trans. Knowl. Discov. Data2
2023 Prompt Tuning Pushes Farther, Contrastive Learning Pulls Closer: A Two-Stage Approach to Mitigate Social Biases
abstract
As the representation capability of Pre-trained Language Models (PLMs) improve, there is growing concern that they will inherit social biases from unprocessed corpora.Most previous debiasing techniques used Counterfactual Data Augmentation (CDA) to balance the training corpus.However, CDA slightly modifies the original corpus, limiting the representation distance between different demographic groups to a narrow range.As a result, the debiasing model easily fits the differences between counterfactual pairs, which affects its debiasing performance with limited text resources.In this paper, we propose an adversarial training-inspired two-stage debiasing model using Contrastive learning with Continuous Prompt Augmentation (named CCPA) to mitigate social biases in PLMs' encoding.In the first stage, we propose a data augmentation method based on continuous prompt tuning to push farther the representation distance between sample pairs along different demographic groups.In the second stage, we utilize contrastive learning to pull closer the representation distance between the augmented sample pairs and then fine-tune PLMs' parameters to get debiased encoding.Our approach guides the model to achieve stronger debiasing performance by adding difficulty to the training process.Extensive experiments show that CCPA outperforms baselines in terms of debiasing performance.Meanwhile, experimental results on the GLUE benchmark show that CCPA retains the language modeling capability of PLMs.
Yingji Li, Mengnan Du, Xin Wang 0035, Ying Wang 0009
ACL (1)1
2023 Measuring and mitigating language model biases in abusive language detection
Rui Song 0008, Fausto Giunchiglia, Yingji Li, Lida Shi, Hao Xu 0012
Inf. Process. Manag.3
2023 Learning continuous dynamic network representation with transformer-based temporal graph neural network
Yingji Li, Mingchen Sun, Ying Wang 0009
Inf. Sci.1
2023 Structural-aware motif-based prompt tuning for graph clustering
Mingchen Sun, Mengduo Yang, Yingji Li, Dongmei Mu, Xin Wang 0035, Ying Wang 0009
Inf. Sci.3