Chunli Xie

dblp:76/9789 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-1490-2991ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MIEM-CA: A Multigranularity Information-Enhanced Entity Matching Method Based on Collaborative Agents
abstract
Entity matching (EM) aims to identify records from different data sources referring to the same real‐world entity. Despite remarkable advances with pretrained language models (PLMs), existing PLM‐based matchers still encounter significant challenges in effectively integrating external knowledge, representing semantic information at multiple granularities, and handling numerical snippets. To address these challenges, we propose a multigranularity information‐enhanced EM method based on collaborative agents (MIEM‐CA), featuring three key components: (1) a multiagent information enhancement module (MI) that leverages extensive external knowledge, the decision‐making and collaboration capabilities of autonomous agents, and the semantic comprehension power of large language models (LLMs), by integrating attribute selection, web search, and feature extraction agents to improve the completeness of entity representation; (2) a multigranularity semantic encoder (ME) that incrementally captures and integrates token‐, attribute‐, and entity‐level semantics, along with their cross‐level correlations, across hierarchical representations spanning the token, attribute, and entity layers (ELs); and (3) a numerical‐aware agent module (NA) that employs the chain‐of‐thought (CoT) strategy to extract numerical information effectively, leverages LLMs to infer the semantic types of these numerical values, and calculates their semantic‐aware numerical similarity. Comprehensive experiments on 10 benchmark datasets, which cover structured, dirty, and textual EM settings, demonstrate that, compared with five baseline methods, MIEM‐CA achieves an average F 1 score improvement of 6.35% on structured datasets, 9.07% on the dirty datasets, and 8.11% across all datasets.
Yaoli Xu, Zheran Yang, Shuaixi Liang, Yongwen Liu, Chunli Xie, Haojie Zhai, Xiayang Shi
Int. J. Intell. Syst.5
2026 Code clone detection based on semantic images
Zexuan Wan, Chunli Xie, You Zeng, Yuanfa Hu
Inf. Softw. Technol.2
2025 FRCS-LLM: A Framework for Refining Code Summarization in Large Language Models via Pre-trained Models
abstract
To enhance software development and maintenance efficiency, code summarization has become a research hotspot. In recent years, due to the generation capabilities of Large Language Models (LLMs), researchers have begun to propose LLMs into code-related tasks. Existing code summarization techniques are insufficient due to their limited ability to deeply comprehend code structure and semantics, frequently yielding verbose or imprecise summaries. To address these challenges, we present FRCS-LLM, a novel code summarization framework that integrates LLMs with pre-trained models to enhance performance and accuracy. Our framework first leverages simulated expert prompts to direct LLMs in generating initial code descriptions, thereby maximizing their linguistic capabilities while minimizing redundancy. We then employ pre-trained models to perform deep fusion of code snippets and functionality descriptions through multimodal joint modeling, extracting both syntactic and semantic information to generate precise, natural, and concise summaries. We evaluate FRCS-LLM on public Java and Python datasets, compared to the strongest baseline, achieving improvements of 2.4%, 3.5%, and 2.0% in BLEU, METEOR, and ROUGE_L metrics on the Java dataset, and improvements of 4.1%, 6.8%, and 3.6% on the Python dataset. Our framework also achieves superior performance in relevance, conciseness, and naturalness, producing summaries that effectively capture code semantics while maintaining completeness, fluency, and grammatical accuracy.
Chunli Xie, Wenbin Zhang 0002
COMPSAC2
2025 SURec: A Semantic-Driven and User Behavior Collaborative Bug Report Recommendation Method
abstract
Bug reports play a crucial role in software development, but the growing volume of reports presents challenges such as duplicate bug reports and information overload. This paper proposes SURec, a method for recommending similar bug reports that integrates semantic analysis with user behavior patterns to improve recommendation accuracy and personalization. SURec extracts semantic features (summaries and descriptions), categorical attributes (products and components), and user behavior features to achieve personalized similarity calculation. Four user information datasets were constructed using Bugzilla (Eclipse, Mozilla) and JIRA (Hadoop, Spark) platforms. Experimental results demonstrate that on large-scale datasets, our method implemented with pre-trained language models achieves higher recommendation accuracy in complex semantic scenarios. In contrast on small-scale datasets with sparse data and noisy categorical attributes, statistical models such as TF-IDF and LSI demonstrate greater robustness. The findings confirm the effectiveness of integrating user historical behavior with multi-dimensional attributes for similar bug report recommendations.
Chunli Xie, Yuanfa Hu, You Zeng
QRS2
2024 Semantic Code Clone Detection Based on Community Detection
abstract
Semantic code clone detection is to find code snippets that are structurally or syntactically different, but semantically identical. It plays an important role in software reuse, code compression. Many existing studies have achieved good performance in non-semantic clone, but semantic clone is still a challenging task. Recently, several works have used tree or graph, such as Abstract Syntax Tree (AST), Control Flow Graph (CFG) or Program Dependency Graph (PDG) to extract semantic information from source codes. In order to reduce the complexity of tree and graph, some studies transform them into node sequences. However, this transformation will lose some semantic information. To address this issue, we propose a novel high-performance method that utilizes community detection to extract features of AST while preserving its semantic information. First, based on the AST of source code, we exploit community detection to split AST into different subtrees to extract the underlying semantics information of different code blocks, and use centrality analysis to quantify the semantic information as the weight of AST nodes. Then, the AST is converted into a sequence of tokens with weights, and a Siamese neural network model is used to detect the similarity of token sequences for semantic code clone detection. Finally, to evaluate our approach, we conduct experiments on two standard benchmark datasets, Google Code Jam (GCJ) and BigCloneBench (BCB). Experimental results show that our model outperforms the eight publicly available state-of-the-art methods in detecting code clones. It is five times faster than the tree-based method (ASTNN) in terms of time complexity.
Zexuan Wan, Chunli Xie, Quanrun Lv, Yasheng Fan
Int. J. Softw. Eng. Knowl. Eng.2
2020 Prespecified-Time Cluster Synchronization of Complex Networks via a Smooth Control Approach
abstract
Most existing finite-/fixed-time synchronization control schemes are nonsmooth or discontinuous, and the settling time is estimated with conservatism. It is due to the utilization of signum function or fraction power state feedback. This brief considers the problem of prespecified-time cluster synchronization of complex networks with a smooth control protocol. The synchronization time is independent of any control parameters or any systems' initial conditions, which is actually uniformly prescribed according to task requirements without any estimations. Moreover, the cluster synchronization can maintain after the specified time, and the smooth control input can always keep uniformly bounded in an infinite time interval as well. Finally, one numerical example is provided to illustrate the effectiveness of the proposed protocol and design method.
Xiaoyang Liu 0002, Daniel W. C. Ho, Chunli Xie
IEEE Trans. Cybern.3
2011 Ontology-Based Reliability Evaluation for Web Service
abstract
Reliability has become a major quality metric for Web service. However, current reliability evaluation approaches lack a formal semantic representation and the support of incomplete or uncertain information. We propose a Web service reliability ontology (WSRO) serving as a basis to characterize the knowledge of Web service. And based on WSRO, a mapping to the probability graphical model is constructed. The Web service reliability evaluation results are obtained by the causality reasoning. Some evaluation results reveal that our approach is applicable and effective.
Xifeng Wang, Bixin Li, Chunli Xie
COMPSAC4
2011 A Staged Model for Web Service Reliability
abstract
In SOA (Service-oriented Architecture), web services consist of five steps: service publishing, service discovery, service composition, service binding and service execution. Faults may occur during every step and cause failure of service execution. Traditional architecture-based reliability models are inapplicable to web services. In this paper, a staged reliability model for web services is proposed which divides the model into multistage model based on the faults occurred at every step. This model considers more failure factors than traditional model.
Chunli Xie, Bixin Li, Xifeng Wang
COMPSAC1
2011 A Web Service Reliability Model Based on Birth-Death Process
Chunli Xie, Bixin Li, Xifeng Wang
SEKE1