Yukun Cao

dblp:96/5464 · DBLP profile ↗
← Back
40ranked-venue papers
25as first author
31since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 14 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 7 · 6 first-author · 7 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 DCTR: Dual-Constraint Subgraph Optimization for Knowledge Graph-based Retrieval-Augmented Generation
abstract
Knowledge Graph (KG)-based Retrieval-Augmented Generation (RAG) shifts the contents of retrieval from narrative text to a relational knowledge network, empowering large language models (LLMs) to harness structured relationships between entities. However, conventional KG-RAG approaches are resource-intensive, requiring either query decomposition with multiple LLM rounds or parameterized static knowledge injection to update the model. Although subgraph reasoning aims to address these issues, most current methods are based on heuristic shortest path and multi-hop graph traversal algorithms. The retrieved subgraphs suffer from incompleteness and semantic drift, and neglect the interaction between subgraph and LLMs in terms of fine-grained structural semantics. We propose a dual-constraint subgraph optimization for KG-RAG (DCTR). It improves subgraph retrieval and generates high-quality subgraphs with structural integrity and information salience for LLMs. Specifically, it formulates subgraph generation as a two-stage graph-theoretic constrained optimization problem to create compact and complete pseudo-labels. Since these pseudo-labels are discrete, a smooth approximation is employed to convert them into a differentiable representation, thereby optimizing the retriever to highlight key information while extracting subgraphs. On two benchmark datasets, DCTR significantly enhances subgraph quality, achieving state-of-the-art performance in LLM reasoning.
Yukun Cao, Luobin Huang, Lisheng Wang
AAAI1
2026 KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering
abstract
Multi-modal Large Language Models (MLLMs) for Visual Question Answering (VQA) often suffer from dual limitations: knowledge hallucination and insufficient fine-grained visual perception.Crucially, we identify that commonsense graphs and scene graphs provide precisely complementary solutions to these respective deficiencies by providing rich external knowledge and capturing fine-grained visual details.However, prior works typically treat them in isolation, overlooking their synergistic potential.To bridge this gap, we propose KG-ViP, a unified framework that empowers MLLMs by fusing scene graphs and commonsense graphs.The core of the KG-ViP framework is a novel retrieval-and-fusion pipeline that utilizes the query as a semantic bridge to progressively integrate both graphs, synthesizing a unified structured context that facilitates reliable multi-modal reasoning.Extensive experiments on FVQA 2.0+ and MVQA benchmarks demonstrate that KG-ViP significantly outperforms existing VQA methods.
Ao Ke, Yukun Cao, Xike Xie
ACL (1)3
2026 P2N2S: Bridging the gap between natural and programming languages for code summarization via large language models
Yijia Tang, YaoShen Yu, Bowei Xia, Yukun Cao
World Wide Web (WWW)5
2025 DPCL-Diff:Temporal Knowledge Graph Reasoning Based on Graph Node Diffusion Model with Dual-Domain Periodic Contrastive Learning
abstract
Temporal knowledge graph (TKG) reasoning that infers future missing facts is an essential and challenging task. Predicting future events typically relies on closely related historical facts, yielding more accurate results for repetitive or periodic events. However, for future events with sparse historical interactions, the effectiveness of this method, which focuses on leveraging high-frequency historical information, diminishes. Recently, the capabilities of diffusion models in image generation have opened new opportunities for TKG reasoning. Therefore, we propose a graph node diffusion model with dual-domain periodic contrastive learning (DPCL-Diff). Graph node diffusion model (GNDiff) introduces noise into sparsely related events to simulate new events, generating high-quality data that better conforms to the actual distribution. This generative mechanism significantly enhances the model's ability to reason about new events. Additionally, the dual-domain periodic contrastive learning (DPCL) maps periodic and non-periodic event entities to Poincaré and Euclidean spaces, leveraging their characteristics to distinguish similar periodic events effectively. Experimental results on four public datasets demonstrate that DPCL-Diff significantly outperforms state-of-the-art TKG models in event prediction, demonstrating our approach's effectiveness. This study also investigates the combined effectiveness of GNDiff and DPCL in TKG tasks.
Yukun Cao, Lisheng Wang, Luobin Huang
AAAI1
2025 GraphInsight: Unlocking Insights in Large Language Models for Graph Structure Understanding
abstract
Although Large Language Models (LLMs) have demonstrated potential in processing graphs, they struggle with comprehending graphical structure information through prompts of graph description sequences, especially as the graph size increases.We attribute this challenge to the uneven memory performance of LLMs across different positions in graph description sequences, known as "Positional bias".To address this, we propose GraphInsight, a novel framework aimed at improving LLMs' comprehension of both macro-and micro-level graphical information.GraphInsight is grounded in two key strategies: 1) placing critical graphical information in positions where LLMs exhibit stronger memory performance, and 2) investigating a lightweight external knowledge base for regions with weaker memory performance, inspired by retrieval-augmented generation (RAG).Moreover, GraphInsight explores integrating these two strategies into LLM agent processes for composite graph tasks that require multi-step reasoning.Extensive empirical studies on benchmarks with a wide range of evaluation tasks show that GraphInsight significantly outperforms all other graph description methods (e.g., prompting techniques and reordering strategies) in understanding graph structures of varying sizes.
Yukun Cao, Zengyi Gao, Zezhong Ding 0001, Xike Xie, Shaohua Kevin Zhou
ACL (1)1
2025 Reinforced and Integrated Prompt Optimization Strategy for Emotion Recognition in Conversation
abstract
Emotion recognition in conversation (ERC) aims to identify the emotion expressed in each utterance within a multi-turn dialogue. In recent years, the widespread adoption of language models (LM) has spurred the development of various prompting paradigms as effective adaptation strategies, aligning LM training objectives with the specific requirements of ERC. However, while the hard prompt-based paradigm offers high human interpretability, it suffers from limited task adaptability due to its discrete nature. In contrast, the soft prompt-based paradigm sacrifices interpretability in favor of improved adaptability by optimizing continuous embedding vectors. Both paradigms generally adopt an invariant prompt across utterances, which restricts their ability to model contextual diversity and results in suboptimal adaptability of the instance. To address these issues, we propose a comprehensive prompting strategy that balances interpretability and adaptability, consisting of two components: reinforced prompt exploration for hard prompts and feature integration for soft prompts. In reinforced prompt exploration, a policy network is trained via reinforcement learning to explore the discrete prompt space under cold-start conditions, efficiently optimizing hard prompts and improving task-level adaptability while preserving interpretability. In feature integration for soft prompts, we incorporate rich semantic features to form contextually relevant soft prompts, assigning each utterance a distinct offset subspace to improve instance-level adaptability. Experiments on three datasets demonstrate that our prompting strategy achieves state-of-the-art performance in ERC.
Yukun Cao, Luobin Huang, Lisheng Wang
ECAI1
2025 Lego Sketch: A Scalable Memory-augmented Neural Network for Sketching Data Streams
abstract
Sketches, probabilistic structures for estimating item frequencies in infinite data streams with limited space, are widely used across various domains. Recent studies have shifted the focus from handcrafted sketches to neural sketches, leveraging memory-augmented neural networks (MANNs) to enhance the streaming compression capabilities and achieve better space-accuracy trade-offs. However, existing neural sketches struggle to scale across different data domains and space budgets due to inflexible MANN configurations. In this paper, we introduce a scalable MANN architecture that brings to life the Lego sketch, a novel sketch with superior scalability and accuracy. Much like assembling creations with modular Lego bricks, the Lego sketch dynamically coordinates multiple memory bricks to adapt to various space budgets and diverse data domains. Theoretical analysis and empirical studies demonstrate its scalability and superior space-accuracy trade-offs, outperforming existing handcrafted and neural sketches.
Yukun Cao, Hairu Wang 0002, Xike Xie, Shaohua Kevin Zhou
ICML2
2025 GCLP: Generative Contrastive Learning with Adaptive Prompt-Guided Diffusion for Temporal Reasoning over Service Knowledge Graphs
Yukun Cao, Lisheng Wang, Luobin Huang
ICSOC (1)1
2025 Prototype-based Optimal Transport for Out-of-Distribution Detection
abstract
Detecting Out-of-Distribution (OOD) inputs is crucial for improving the reliability of deep neural networks in the real-world deployment. In this paper, inspired by the inherent distribution shift between in-distribution (ID) and OOD data, we propose a novel method that leverages optimal transport to measure the distribution discrepancy between test inputs and ID prototypes. The resulting transport costs are used to quantify the individual contribution of each test input to the overall discrepancy, serving as a desirable measure for OOD detection. To address the issue that solely relying on the transport costs to ID prototypes is inadequate for identifying OOD inputs closer to ID data, we generate virtual outliers to approximate the OOD region via linear extrapolation. By combining the transport costs to ID prototypes with the costs to virtual outliers, the detection of OOD data near ID data is emphasized, thereby enhancing the distinction between ID and OOD inputs. Extensive evaluations demonstrate the superiority of our method over state-of-the-art methods.
Ao Ke, Chuanwen Feng, Yukun Cao, Xike Xie, Shaohua Kevin Zhou, Lei Feng 0006
IJCAI4
2025 DIPF: Dual-Intervention Prompt Framework for Emotion Recognition in Conversation
abstract
Emotion Recognition in Conversation (ERC) is a challenging task in conversational AI systems. However, current ERC methods that rely on language models (LM) often use fixed prompts, which lack dynamic adaptability and fail to adequately incorporate external knowledge influencing conversational emotions. To address these limitations, this paper proposes the Dual-Intervention Prompt Framework for emotion recognition in conversation (DIPF). DIPF introduces two types of prompts: explicit prompts and implicit prompts. Explicit prompts leverage external knowledge via knowledge graphs, allowing the model to better simulate conversational flow and emotional dynamics. Implicit prompts are integrated into the hidden layers of the LM, becoming part of the model’s representation and participating in the gradient descent process. Additionally, DIPF manipulates the model representations corresponding to implicit prompts using a low-rank projection matrix, guiding the model’s behavior during inference and improving the optimization of implicit prompts. Extensive experiments on three benchmark datasets demonstrate that DIPF achieves relative F1 score improvements of 1.27% on IEMOCAP, 2.26% on MELD, and 0.72% on EmoryNLP over the strongest baseline (EACL), with an average relative improvement of 1.52%. These results confirm the effectiveness of DIPF in enhancing emotional understanding through dynamic and knowledge-aware prompting strategies.
Yukun Cao, Luobin Huang, Lisheng Wang
IJCNN1
2025 Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
abstract
Large Language Models have excelled in various domains but face efficiency challenges due to the growing Key-Value (KV) cache required for long-sequence inference. Recent efforts aim to reduce KV cache size by evicting vast non-critical cache elements during runtime while preserving generation quality. However, these methods typically allocate compression budgets uniformly across all attention heads, ignoring the unique attention patterns of each head. In this paper, we establish a theoretical loss upper bound between pre- and post-eviction attention output, explaining the optimization target of prior cache eviction methods, while guiding the optimization of adaptive budget allocation. Base on this, we propose {\it Ada-KV}, the first head-wise adaptive budget allocation strategy. It offers plug-and-play benefits, enabling seamless integration with prior cache eviction methods. Extensive evaluations on 13 datasets from Ruler and 16 datasets from LongBench, all conducted under both question-aware and question-agnostic scenarios, demonstrate substantial quality improvements over existing methods. Our code is available at https://github.com/FFY0/AdaKV.
Junlin Lv, Yukun Cao, Xike Xie, Shaohua Kevin Zhou
NeurIPS3
2025 Gradient Reweighting-Based Representation Intervention and Prompting Framework for Emotion Recognition in Conversation
Yukun Cao, Lisheng Wang, Luobin Huang, Yongcheng He
PRICAI1
2025 DMCM: Dual-space Mapping Contrastive Meta Learning for Cold Start Recommendation*
abstract
In recent years, meta-learning-based and contrastive learning-based approaches have achieved promising results in addressing the cold-start problem in recommendation systems by globally modeling user preferences. However, due to the sparse historical interactions between cold-start items and users, it is difficult to accurately capture the hierarchical relationships between users and items. To alleviate the above issue, we propose a dual-space mapping contrastive meta learning for cold start recommendation (DMCM), which simultaneously maps user and item information into both poincaré space and euclidean space to more effectively capture the latent hierarchical structure of user-item interaction information. Through dual-space mapping contrastive learning, the representational capacity of user-item associations is enhanced, and the coverage of user interaction information is expanded. Subsequently, a multi-level meta-learning path dynamic adjustment strategy is employed to further improve the model ability to personalize and adapt to new users and new items. Experiments conducted on three public datasets demonstrate that the proposed framework outperforms baseline methods in both cold-start and non-cold-start scenarios, significantly improving recommendation accuracy.
Yukun Cao, Niu Gu, Yongcheng He
SMC1
2025 CDCO: Cross-Domain Contrastive Optimization Framework for Enhancing Multi-Task Learning in Small Pre-trained Language Models
abstract
The adapter technique has significantly improved the performance of fine-tuning pre-trained language models (PLMs) in multi-task settings, especially for small models with limited resources. However, current adapter frameworks typically rely on single-task data, a fixed representation space, and a single training phase. This approach limits their ability to generalize across domains in multi-task scenarios, thus restricting performance improvements for smaller models. To address these challenges, we propose the cross-domain contrastive optimization (CDCO) framework to enhance performance in multi-task learning. CDCO improves model performance by asynchronously co-optimizing across diverse task data sources, representation spaces, and multi-stage structures. Specifically, CDCO introduces innovations in both data sample selection and training strategy. First, CDCO introduces out-of-domain manifold sampling (ODMS), which enhances training diversity by selecting challenging hard-negative samples from out-of-domain datasets through manifold learning. Second, CDCO employs multi-stage asynchronous co-optimization (MAC), mapping samples from ODMS to Euclidean and Poincaré spaces. Then, it constructs a cross-domain contrastive loss based on the spatial properties of these distributions to guide the optimization process. By sequentially optimizing adapter layers across different spatial distributions, CDCO maximizes the potential of the adapter while mitigating overfitting, thus improving the adaptability and stability of small models in multi-task environments. Experimental results demonstrate that CDCO significantly improves performance on in-domain (ID), out-of-domain (OOD), and knowledge-intensive (KI) tasks, confirming its broad applicability and effectiveness.
Yukun Cao, Yongcheng He, Niu Gu
SMC1
2025 DENL: Dynamic Emotion Neural Link for Efficient Emotion Recognition in Conversations *
abstract
The Emotion Recognition in Conversations (ERC) task requires models to precisely capture subtle emotional nuances within contextual environments. Presently, the correlation between utterances and emotions is relatively weak, meaning the same utterance might express entirely different emotions. To tackle this challenge, pre-trained language models (PLMs) are typically employed, either through full-parameter fine-tuning or parameter-efficient tuning and learning (PETL) methods. However, these methods incur high computational costs. To address the weak correlation between utterances and emotions under limited computational resources, we propose a feature-task dynamic emotion neural link (DENL), which refines emotional feature representation at a low computational cost and rapidly adapts to ERC tasks. At the feature-learning layer, we embed multiple specialized Emotion Dual-Rank Adapters (EDRA) in parallel, coupled with a gradient-aware dynamic gating mechanism (DeepGate), to avoid propagation costs through the backbone network during backpropagation, thus creating a low-cost emotion neural link to capture emotional features. At the task-learning layer, we utilize a Emotion Space Intervention (ESI) approach, employing a low-rank subspace to manipulate portions of the hidden representations of utterance embeddings, thus guiding the model to rapidly adapt to ERC tasks. Experimental results demonstrate that DENL not only improves the accuracy of fine-grained emotion classification on three benchmark ERC datasets, but also significantly reduces computational cost and parameter size.
Yukun Cao, Yongcheng He, Niu Gu
SMC1
2025 LEGO-GraphRAG: Modularizing Graph-based Retrieval-Augmented Generation for Design Space Exploration
abstract
GraphRAG integrates (knowledge) graphs with large language models (LLMs) to improve reasoning accuracy and contextual relevance. Despite its promising applications and strong relevance to multiple research communities, such as databases and natural language processing, GraphRAG currently lacks modular workflow analysis, systematic solution frameworks, and insightful empirical studies. To bridge these gaps, we propose LEGO-GraphRAG , a modular framework that enables: 1 ) fine-grained decomposition of the GraphRAG workflow, 2 ) systematic classification of existing techniques and implemented GraphRAG instances, and 3 ) creation of new GraphRAG instances. Our framework facilitates comprehensive empirical studies of GraphRAG on large-scale real-world graphs and diverse query sets, revealing insights into balancing reasoning quality, runtime efficiency, and token or GPU cost, that are essential for building advanced GraphRAG systems.
Yukun Cao, Zengyi Gao, Xike Xie, Shaohua Kevin Zhou, Jianliang Xu
Proc. VLDB Endow.1
2025 PeTracker: Poincaré-Based Dual-Strategy Emotion Tracker for Emotion Recognition in Conversation
abstract
With the increasing use of interactive applications, the importance of Emotion recognition in conversation (ERC) is growing. Current research in the ERC domain mainly emphasizes the extraction of contextual information. However, challenges arise due to multi-turn conversation scenarios and the natural transformation of emotions, particularly in identifying subtle emotion transfers. Moreover, emotions exhibit nonlinear characteristics in semantic spaces, leading to potential confusion when discerning similar semantic emotions in the Euclidean semantic space. To address these issues, this study proposes a Poincaré-based dual-strategy emotion tracker for emotion recognition in conversation (PeTracker), which introduces the hyperbolic space representation in the ERC domain. Based on the spatial properties of the hyperbolic space representation to capture the nonlinear relationships among features, PeTracker encompasses two learning strategies. Poincaré emotional geometry curriculum learning (PGCL) and Poincaré emotional stratification contrastive learning (PSCL). In PGCL, the similarity of emotion labels is effectively discerned using the Poincaré distance, quantifying emotion transfer distances and facilitating the identification of subtle emotion transfers in utterance. In PSCL, PeTracker extracts and adapts multi-level features, mapping them to the Poincaré ball space to build emotion prototype-based contrastive learning. This process enhances the model’s ability to distinguish between similar emotion labels. while alleviating potential label confusion issues. Experimental results on three general datasets demonstrate that PeTracker achieves optimal or near-optimal performance. Furthermore, the study investigates the role and impact of the Poincaré ball in differentiating similar emotions.
Yukun Cao, Luobin Huang, Yijia Tang
IEEE Trans. Affect. Comput.1
2024 FRELinker: A Novel Issue-Commit Link Recovery Model Based on Feature Refinement and Expansion with Multi-Classifier Fusion
abstract
In the field of software traceability (ST), machine learning (ML) has become a common and effective method for automated issue-commit link recovery. The features extracted from issue and commit artifacts are composed of significantly different types of data, such as issue summaries, diff codes, and hashes. Such complex and diverse data poses a challenge to conventional ML methods. To overcome this challenge, we propose a novel model named FRELinker, which trains independent classifiers based on the type of issue-commit training data, fully leveraging the effectiveness of ML for single type data, and then fuses the classifiers. Specifically, we categorize the features into four types: textual features, code features, non-textual features, and similarity features, and extend text similarity features by adding hybrid textual similarity measures. And then, we use a ranking method to select the optimal classifiers corresponding to these four types of features. Among them, the optimal classifier for textual features is Gradient Boosting (GB), the optimal classifier for code features is Logistic Regression (LR), and the optimal classifier for non-textual features and similarity features is Random Forest (RF). Finally, we use a Bayesian optimization model to fuse these four classifiers. Experimental results show that our method outperforms competing methods Hybrid-Linker and DeepLink in terms of Precision, Recall, and F-measure on six real-world open-source software (OSS) datasets, demonstrating significant performance advantages in complex and diverse data.
Bangchao Wang, Hongyan Wan, Jiaxu Zhu, Yukun Cao
APSEC6
2024 HANTracer: Leveraging Heterogeneous Graph Attention Network for Large-Scale Requirements-Code Traceability Link Recovery
abstract
In the task of requirements-to-code traceability link recovery, the continues growth of software scale has led to diminishing differences between indices and more complex nonlinear relationships within the data. This results in the performance decline of the most widely used information retrieval methods and machine learning methods when handling this task. Therefore, we propose a requirement traceability method based on heterogeneous graph attention networks, named HANTracer. The model integrates the high-dimensional vectors generated by a pre-trained model as node features to deepen the dif-ferentiation between nodes and enhances node representations with contextual information learned from the graph structure. Additionally, it utilizes the properties of heterogeneous edges to construct various edge features, such as code calling, code inheritance, and text similarity relationships, aiding the model in understanding and utilizing the relationships between different types of data elements. By incorporating average pooling layers and multiple fully connected layers, the HANTracer model is further improved to enhance its ability to extract nonlinear features. Experimental results indicate that HANTracer achieves an average F1 performance higher than the state-of-the-art methods TAROT by 100.30% and DF4RT by 46.27% on seven real-world open source software (OSS) datasets, demonstrating significant performance advantages in large-scale and complex data environments.
Zhiyuan Zou, Bangchao Wang, Hongyan Wan, Huan Jin, Yukun Cao
APSEC6
2024 Mayfly: a Neural Data Structure for Graph Stream Summarization
abstract
A graph is a structure made up of vertices and edges used to represent complex relationships between entities, while a graph stream is a continuous flow of graph updates that convey evolving relationships between entities. The massive volume and high dynamism of graph streams promote research on data structures of graph summarization, which provides a concise and approximate view of graph streams with sub-linear space and linear construction time, enabling real-time graph analytics in various domains, such as social networking, financing, and cybersecurity. In this work, we propose the Mayfly, the first neural data structure for summarizing graph streams. The Mayfly replaces handcrafted data structures with better accuracy and adaptivity. To cater to practical applications, Mayfly incorporates two offline training phases. During the larval phase, the Mayfly learns basic summarization abilities from automatically and synthetically constituted meta-tasks, and in the metamorphosis phase, it rapidly adapts to real graph streams via meta-tasks. With specific configurations of information pathways, the Mayfly enables flexible support for miscellaneous graph queries, including edge, node, and connectivity queries. Extensive empirical studies show that the Mayfly significantly outperforms its handcrafted competitors.
Yukun Cao, Hairu Wang 0002, Xike Xie, Shaohua Kevin Zhou
ICLR2
2024 FRD-DST: A Fine-Grained Relation-Driven Model for Dialogue State Tracking
abstract
Dialogue State Tracking (DST) aims to convert dialogue history into dialogue states of slot-value pairs. Many existing studies usually utilize deep neural networks to learn the representation of dialogues and slots. However, these studies usually do not adequately consider the fine-grained relation-ships between each word and slot in a dialogue. Meanwhile, due to the complexity of the task, there needs to be more textual and semantic diversity in dialogues. To address these challenges, we propose a fine-grained relation-driven model (FRD-DST) to synthesize the dialogues and the connections between words and slots. In this approach, each word and slot of information in a dialogue context is constructed as a dialogue word-slot heterograph, and a relationship aggregation network is used to capture the fine-grained features between them. Meanwhile, to complement the contextual association features that the relational aggregation network may not adequately capture, we use a conditional random field (CRF) to capture the dialogue contextual association features after synonym replacement. The two feature information sets are fused in a hidden space, which can create new features by mixing the hidden states of different texts to create new semantic variants, thus enhancing the diversity of dialogue texts and semantics. The experimental results show that FRD-DST achieves state-of-the-art DST performance on the MultiWOZ 2.1 and MultiWOZ 2.4 datasets compared to existing DST methods.
Yukun Cao, Yuanmin Liu
SMC1
2024 NAEP: Neighborhood Multihead Attention with Evidential Conditional Neural Processes for Few-Shot Knowledge Graph Completion
abstract
In few-shot knowledge graph completion (FKGC), inferring new relationships from limited data is a key challenge. Current FKGC frameworks often struggle to detect subtle differences among entity types and handle noise in sparse environments, impacting their performance. Aligning relational representations with entity and relationship data, while managing prediction uncertainties, is another significant hurdle. We propose the NAEP model, which integrates a hierarchical relational gating graph attention network to discern complex relational dynamics and differentiate entity types. It employs neighborhood multihead attention for path encoding to reduce noise and incorporates an Evidential Conditional Neural Processes encoder to handle predictive uncertainties. Preliminary results show that NAEP performs robustly across various datasets, from large ones like NELL and WIKI to smaller ones like FB15K-237, demonstrating its versatility and superior performance.
Yukun Cao, Luobin Huang, Yuanmin Liu
SMC1
2024 XWCoDe: XGBoost with Weighted Code Dependency for Requirements-to-Code Traceability Link Recovery
abstract
Information Retrieval (IR), Machine Learning (ML), and Deep Learning (DL) have become mainstream methods for traceability link recovery. However, IR-based methods face the challenge of low precision, while DL-based methods require large-scale training data to achieve better performance. In this paper, we propose a novel model XWCoDe, which apply XGBoost combined with a weighted code dependency strategy to traceability link recovery domain. In order to refine the initial candidate links generated by the XGBoost model, the strategy only modifies low confidence candidate links and pioneers the use of graph embedding technology node2vec to calculate the importance of each code dependency relationship. The experimental results show that the average F1 score of XW CoDe on 4 datasets and 9 training/testing ratios is 12.93 % higher than the state-of-the-art method DF4RT.
Zhiyuan Zou, Bangchao Wang, Hongyan Wan, Zhiquan An, Yukun Cao
SMC6
2024 Learning to Sketch: A Neural Approach to Item Frequency Estimation in Streaming Data
abstract
Recently, there has been a trend of designing neural data structures to go beyond handcrafted data structures by leveraging patterns of data distributions for better accuracy and adaptivity. Sketches are widely used data structures in real-time web analysis, network monitoring, and self-driving to estimate item frequencies of data streams within limited space. However, existing sketches have not fully exploited the patterns of the data stream distributions, making it challenging to tightly couple them with neural networks that excel at memorizing pattern information. Starting from the premise, we envision a pure neural data structure as a base sketch, which we term the meta-sketch, to reinvent the base structure of conventional sketches. The meta-sketch learns basic sketching abilities from meta-tasks constituted with synthetic datasets following Zipf distributions in the pre-training phase and can be quickly adapted to real (skewed) distributions in the adaption phase. The meta-sketch not only surpasses its competitors in sketching conventional data streams but also holds good potential in supporting more complex streaming data, such as multimedia and graph stream scenarios. Extensive experiments demonstrate the superiority of the meta-sketch and offer insights into its working mechanism.
Yukun Cao, Hairu Wang 0002, Xike Xie, Shaohua Kevin Zhou
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Word Sense Disambiguation Combining Knowledge Graph and Text Hierarchical Structure
abstract
Current supervised word sense disambiguation models have obtained high disambiguation results using annotated information of different word senses and pre-trained language models. However, the semantic data of the supervised word sense disambiguation models are in the form of short texts, and much of the corpus information is not rich enough to distinguish the semantics in different scenarios. This article proposes a bi-encoder word sense disambiguation method combining a knowledge graph and text hierarchy structure, by introducing structured knowledge from the knowledge graph to supplement more extended semantic information, using the hierarchy of contextual input text to describe the meaning of words and phrases, and constructing a BERT-based bi-encoder, introducing a graph attention network to reduce the noise information in the contextual input text, so as to improve the disambiguation accuracy of the target words in phrase form and ultimately improve the disambiguation effectiveness of the method. By comparing the method with the latest nine comparison algorithms in five test datasets, the disambiguation accuracy of the method mostly outperformed the comparison algorithms and achieved better results.
Yukun Cao, ChengKun Jin, Yijia Tang, Ziyue Wei
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2023 Meta-Sketch: A Neural Data Structure for Estimating Item Frequencies of Data Streams
abstract
To estimate item frequencies of data streams with limited space, sketches are widely used in real applications, including real-time web analytics, network monitoring, and self-driving. Sketches can be viewed as a model which maps the identifier of a stream item to the corresponding frequency domain. Starting from the premise, we envision a neural data structure, which we term the meta-sketch, to go beyond the basic structure of conventional sketches. The meta-sketch learns basic sketching abilities from meta-tasks constituted with synthetic datasets following Zipf distributions in the pre-training phase, and can be fast adapted to real (skewed) distributions in the adaption phase. Extensive experiments demonstrate the performance gains of the meta-sketch and offer insights into our proposals.
Yukun Cao, Xike Xie
AAAI1
2023 Aware-Transformer: A Novel Pure Transformer-Based Model for Remote Sensing Image Captioning
Yukun Cao, Jialuo Yan, Yijia Tang, Zhenyi He, Kangle Xu
CGI (1)1
2023 Learn to Explore: on Bootstrapping Interactive Data Exploration with Meta-learning
abstract
Interactive data exploration (IDE) is an effective way of comprehending big data, whose volume and complexity are beyond human abilities. The main goal of IDE is to discover user interest regions from a database through multi-rounds of user labelling. Existing IDEs adopt active-learning framework, where users iteratively discriminate or label the interestingness of selected tuples. The process of data exploration can be viewed as the process of training of a classifier, which determines whether a database tuple is interesting to a user. An efficient exploration thus takes very few iterations of user labelling to reach the data region of interest. In this work, we consider the data exploration as the process of few-shot learning, where the classifier is learned with only a few training examples, or exploration iterations. To this end, we propose a learning-to-explore framework, based on meta-learning, which learns how to learn a classifier with automatically generated meta-tasks, so that the exploration process can be much shortened. Extensive experiments on real datasets show that our proposal outperforms existing explore-by-example solutions in terms of accuracy and efficiency.
Yukun Cao, Xike Xie
ICDE1
2023 Dialogue-Clues: Dual-channel Dialogue Clues Embedding Context Perception Network for Emotion Recognition in Conversations
abstract
In recent years, emotion recognition in conversations has gained widespread attention due to its extensive applications. Many recent studies have focused on perceiving conversational context from the perspective of capturing dialogue clues. However, these studies often use the entire conversation sequence to represent the dialogue clues, which may result in insufficient representation of the speaker's emotional dynamics in multi-turn conversation. To address these issues, we propose a novel dual-channel dialogue clues embedding context perception network (Dialogue-Clues) that integrates dialogue clues information into global conversation context modeling. We also introduce a dual-channel dialogue clues perception architecture that captures and reinforces both static and dynamic dialogue clues in conversations, and bidirectionally reinforces both types of dialogue clues information. To better represent dynamic dialogue clues, we construct a novel speaker-interaction isomorphic graph structure for the dual-channel dialogue clues perception architecture. Through extensive comparisons with ten existing methods on four public datasets, we confirm the effectiveness of the proposed method. The results demonstrate that integrating Dialogue-Clues information can improve the ability of context modeling.
Yukun Cao, Zhenyi He, Yijia Tang, Jialuo Yan, Kangle Xu
SMC1
2023 Heterogeneous Reinforcement Learning Network for Aspect-Based Sentiment Classification With External Knowledge
abstract
Aspect-based sentiment classification aims to automatically predict the sentiment polarity of the specific aspect in a text. However, it is challenging to confirm the mapping between the aspect and the core context since a number of existing methods concentrate on building the global relations of the full context rather than the partial connections based on the aspects. Motivated by the fundamental insights of reinforcement learning, we propose a novelHeterogeneousReinforcementLearningNetwork for aspect-based sentiment analysis (HRLN) to alleviate these issues, which contains two primary components, a heterogeneous network module, and a knowledge graph-based reinforcement learning module consistent with common-sense knowledge and emotional knowledge. To evaluate the effectiveness of HRLN, we conduct extensive experiments on five benchmark datasets, which indicate that HRLN achieves competitive performance and yields state-of-the-art results on all datasets. Additionally, we present an intuitive comprehension of why our HRLN model is more robust for aspect-based sentiment classification via case studies.
Yukun Cao, Yijia Tang, Haizhou Du, Ziyue Wei, ChengKun Jin
IEEE Trans. Affect. Comput.1
2022 Isomer: Transfer enhanced dual-channel heterogeneous dependency attention network for aspect-based sentiment classification
Yukun Cao, Yijia Tang, Haizhou Du, Ziyue Wei, ChengKun Jin
Knowl. Based Syst.1
2020 Learning Directional Feature Maps for Cardiac MRI Segmentation
Yukang Wang, Heshui Shi, Yukun Cao, Dandan Tu, Changzheng Zhang, Yongchao Xu
MICCAI (4)5
2020 Pay More Attention to Discontinuity for Medical Image Segmentation
Jiajia Chu, Yajie Chen, Heshui Shi, Yukun Cao, Dandan Tu, Richu Jin, Yongchao Xu
MICCAI (4)5
2019 Exploring Semantic Change of Chinese Word Using Crawled Web Data
Xiaofei Xu 0002, Yukun Cao, Li Li 0006
ICWE2
2019 Social-Aware and Sequential Embedding for Cold-Start Recommendation
Yukun Cao, Li Li 0006, Li Liu 0001, Jun Liao 0001
KSEM (1)2
2019 DST: A Deep Urban Traffic Flow Prediction Framework Based on Spatial-Temporal Features
Yukun Cao, Li Li 0006
KSEM (1)2
2007 An intelligent fuzzy-based recommendation system for consumer electronic products
Yukun Cao
Expert Syst. Appl.1
2005 A Novel Framework for Web Page Classification Using Two-Stage Neural Network
Yukun Cao, Qingsheng Zhu
ADMA2
2005 An On-line Intelligent Recommendation System for Digital Products Using Fuzzy Logic
Yukun Cao, Chengliang Wang 0002
WISE1
2004 An E-mail Filtering Approach Using Neural Network
Yukun Cao, Xiaofeng Liao 0001
ISNN (2)1