EDBT 2026 Demo / reviewers in the wild / expert
Xuan Zhang 0002
dblp:36/31-2
· DBLP profile ↗
44ranked-venue papers
2as first author
41since 2021 · last 2027
0000-0003-2929-2126ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 18 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Computer networks · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Diffusion sequential recommendation model based on item co-occurrence anchoring and user semantic collaboration
Shaoheng Xie, Xuan Zhang 0002, Kunpeng Du, Jishu Wang, Junda Li, Weiyi Shang, Zhi Jin 0001 |
Expert Syst. Appl. | 2 |
| 2026 | World of Logs: A Dataset of Logs from Online DocumentsabstractSoftware logs serve as valuable resources for understanding system running and are extensively used in diverse software maintenance tasks. As software systems get more complex and log data grows, a good log dataset is fundamental for developing automated log analysis tools. However, current log datasets are limited in three aspects, i.e., narrow in scope, lacking context information, and outdated. To bridge this gap, in this paper, we aim to extract software logs from online resources (e.g., JIRA issue reports, GitHub repositories, and Stack Overflow discussions), which concern various types of software systems and provide context for logs, such as observed behaviors and expected behaviors. This work introduces WoL, a dataset comprising real-world logs along with their contextual information. WoL currently contains over 2.5 million log messages or logging statements from diverse online resources and is publicly available to facilitate reproducible research. WoL can be used for various log-related tasks, including understanding logging intent and quality, anomaly detection, and linking logs with software artifacts for contextual analysis. WoL is publicly available on Zenodo and will be continuously updated. Furthermore, based on WoL, we develop a search engine, LogSearch, to support user queries. Kundi Yao, Lizhi Liao, Pengyu Nie 0001, Xuan Zhang 0002, Weiyi Shang |
MSR | 5 |
| 2026 | Learning to Evolve: Bayesian-Guided Continual Knowledge Graph Embedding
LinYu Li 0001, Zhi Jin 0001, Yuanpeng He, Dongming Jin, Yichi Zhang 0009, Haoran Duan 0002, Xuan Zhang 0002, Zhengwei Tao, Nyima Tashi |
WWW | 7 |
| 2026 | A knowledge graph-driven generation framework for perceptual decomposition and serial logical reasoning with large language models
Xuan Zhang 0002, Kunpeng Du, Junda Li, LinYu Li 0001, Tong Li 0004, Zhi Jin 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Collaborative denoising and semantic preservation: A sequential recommendation model via frequency-aware diffusion and KAN
Jiangbin Chen, Xuan Zhang 0002, Kunpeng Du, Shaoheng Xie, Rui Zhu 0009, Junda Li, Weiyi Shang, Zhi Jin 0001 |
Expert Syst. Appl. | 2 |
| 2026 | Using external knowledge to enhance user preferences for better sequential recommendation
Yubin Ma, Xuan Zhang 0002, Zhi Jin 0001, Weiyi Shang, Chen Gao 0006, LinYu Li 0001 |
Expert Syst. Appl. | 3 |
| 2026 | BTCP: A Blockchain-Based Trusted Data Sharing Framework With Congestion Control and Proximity Evaluation for Internet of VehiclesabstractInternet of Vehicles (IoV) communication serves as a critical enabler for intelligent transportation systems, where the reliability of data sharing fundamentally determines overall system performance. Data reliability not only dictates the quality of IoV services but also constitutes a core factor in ensuring traffic safety and improving road operational efficiency. At present, IoV data sharing confronts several key challenges: (i) degradation of message credibility due to malicious attackers; (ii) ineffectiveness of conventional reputation mechanisms under high vehicular mobility; and (iii) channel congestion and information loss caused by redundant data transmissions. To address these issues, this paper proposes a Hybrid Hash Chord Protocol, which integrates Geohash geocoding with the Chord distributed lookup algorithm to establish a location-aware peer-to-peer data forwarding mechanism. This approach enables efficient data transmission with reduced hop count. In addition, a recursive traffic data filtering method is introduced to effectively suppress duplicate data reporting. Moreover, a Bayesian Dynamic Fading Reputation Model is developed, incorporating time decay factors and historical behavior weighting to mitigate intelligent attacks and node inertia, thereby enhancing road safety and traffic efficiency. Experimental results indicate that, within an IoV data-sharing context, the proposed model improves traffic efficiency by approximately 4% and reduces overall communication overhead by 36% compared to existing schemes. When the reputation threshold is set to 0.35, the model achieves a 100% malicious vehicle detection rate, whereas benchmark methods remain below 80% under the same threshold. In summary, the proposed framework achieves a notable balance among enhancing data trustworthiness, optimizing traffic performance, and minimizing communication costs, offering a practicable solution for building efficient and reliable IoV data-sharing systems. Rui Zhu 0009, Zhenyu Xue, Junqiao Song, Abdelsalam Helal, Xuan Zhang 0002, Yeting Chen |
IEEE Internet Things J. | 6 |
| 2026 | IVC-DB: Iterative verification correction method guided by dual-Backward mathematical reasoning in large language models
Kunpeng Du, Xuan Zhang 0002, Chen Gao 0006, Rui Zhu 0009, Tong Li 0004, Zhi Jin 0001 |
Knowl. Based Syst. | 2 |
| 2026 | BPO-CBS: A Data-Driven Blockchain Performance Optimization Framework for Cloud Blockchain ServicesabstractRecently, blockchain has been widely used in important scenarios (e.g., finance and auditing). To fully meet the needs of various business scenarios and reduce deployment costs, cloud blockchain services (CBS) are now being offered by cloud computing providers. However, in high-frequency and large-scale transaction scenarios, blockchain performance faces serious challenges, limiting its further application. Therefore, blockchain performance optimization (BPO) has become a key field. Recent BPO methods that adjust blockchain configuration parameters like block size, offer benefits such as low cost and easy deployment. However, these methods face challenges including unsuitability for dynamic environments, high optimization overhead, and failure to consider marginal utility (MU) in BPO. MU describes the decreasing effectiveness of BPO as transaction arrival rates increases, eventually leading to limited BPO benefits. This paper proposes a data-driven BPO framework (BPO-CBS) for CBS. First, a blockchain performance prediction model is trained using ensemble learning. Second, a performance scoring and adjustment mechanism is designed to identify optimal configuration parameters and adjust them to enhance BPO. Finally, extensive quantitative and qualitative comparisons with related works show that BPO-CBS achieves more effective BPO with low optimization overhead. Jishu Wang, Xuan Zhang 0002, Linfeng Liu 0007, Xuekun Yang, Chen Miao, Rui Zhu 0009, Zhi Jin 0001 |
IEEE Trans. Cloud Comput. | 2 |
| 2026 | Towards Structure-Aware Model for Multi-Modal Knowledge Graph CompletionabstractKnowledge graphs (KGs) play a key role in promoting various multimedia and AI applications. However, with the explosive growth of multi-modal information, traditional knowledge graph completion (KGC) models cannot be directly applied. This has attracted a large number of researchers to study multi-modal knowledge graph completion (MMKGC). Since MMKG extends KG to the visual and textual domains, MMKGC faces two main challenges: (1) how to deal with the fine-grained modality information interaction and awareness; (2) how to ensure the dominant role of graph structure in multi-modal knowledge fusion and deal with the noise generated by other modalities during modality fusion. To address these challenges, this paper proposes a novel MMKGC model named TSAM, which integrates fine-grained modality interaction and dominant graph structure to form a high-performance MMKGC framework. Specifically, to solve the challenges, TSAM proposes the Fine-grained Modality Awareness Fusion method (FgMAF), which uses pre-trained language models better to capture fine-grained semantic information interaction of different modalities and employs an attention mechanism to achieve fine-grained modality awareness and fusion. Additionally, TSAM presents the Structure-aware Contrastive Learning method (SaCL), which utilizes two contrastive learning approaches to align other modalities more closely with the structured modality. Extensive experiments show the proposed TSAM model significantly outperforms existing MMKGC models on widely used multi-modal datasets. The code is available athttps://github.com/2391134843/TSAM. LinYu Li 0001, Zhi Jin 0001, Yichi Zhang 0009, Dongming Jin, Chengfeng Dou, Yuanpeng He, Xuan Zhang 0002, Haiyan Zhao 0001 |
IEEE Trans. Multim. | 7 |
| 2025 | MASFlow: Multi-Agent Based Service Workflow GenerationabstractService workflows are fundamental in software ser-vice systems where standardized processes are implemented as workflow models to achieve automated service execution. Despite the obvious benefits of model-driven architecture, its mainstream adoption is hampered by the need for extensive knowledge and sophisticated modeling abilities to create such models. Recently, multi-agent frameworks have thrived in handling complex tasks by harnessing collective intelligence to tackle intricate problems. MASFlow, a progressive multi-agent collaboration framework for automated service workflow model generation, is presented in this paper. Through the use of specialized agents that represent various team responsibilities, MASFlow replicates real-world design cooperation by breaking down the modeling process into three coordinated phases: Structuring, Orchestration, and Re-view. Experimental results demonstrate that MASFlow effectively mitigates the hallucination generation phenomenon commonly observed in large language models (LLMs) when handling complex service workflows through a phased task decomposition strategy. With an accuracy rate of 92.83%, the generated service workflow model demonstrated notable advantages above existing mainstream neural network architecture techniques and solutions that directly use LLMs. Rui Zhu 0009, Jiapeng Chen, Tianrui Bai, Hua Yue, Jianglong Qin, Xuan Zhang 0002 |
SSE | 6 |
| 2025 | Promoting Unsupervised Data-To-Text Generation Using Retraining and Unified LinearizationabstractABSTRACT In recent years, many studies have focused on unsupervised data‐to‐text generation methods. However, existing unsupervised methods still require a large amount of unlabeled sample training, leading to significant data collection overhead. We propose a low‐resource unsupervised method called CycleRUR. This method first converts various forms of structured data (such as tables, knowledge graph(KG) triples, and meaning representations(MR)) into unified KG triples to improve the model's ability to adapt to different structured data. Additionally, CycleRUR incorporates a retraining module and a contrastive learning module within a cycle training framework, enabling the model to learn and converge from a small amount of unpaired KG triples and reference text corpus, thereby improving the model's accuracy and convergence speed. We evaluated the model's performance on the WebNLG and E2E datasets. Using only 10% of unpaired training data, our method achieved the effects of fully supervised fine‐tuning. On the WebNLG dataset, it resulted in an 18.41% improvement in METEOR compared to supervised models. On the E2E dataset, it achieved improvements of 1.37% in METEOR and 4.97% in BLEU. Experiments also demonstrated that under unified linearization, CycleRUR exhibits good generalization capabilities. Xuan Zhang 0002, Kunpeng Du, Chen Gao 0006, Zhuxian Ma |
Concurr. Comput. Pract. Exp. | 2 |
| 2025 | Knowledge-enhanced prototypical network with graph structure and semantic information interaction for low-shot joint spoken language understanding
Kunpeng Du, Xuan Zhang 0002, Chen Gao 0006, Weiyi Shang, Yubin Ma, Zhi Jin 0001, LinYu Li 0001 |
Expert Syst. Appl. | 2 |
| 2025 | Task-Oriented Dynamic Knowledge Distillation for Continuous Few-Shot Relation Extraction
Hexing Yang, Xuan Zhang 0002, Chen Gao 0006, Weiyi Shang, Kunpeng Du, Tong Li 0004 |
Knowl. Based Syst. | 2 |
| 2025 | Neighborhood structure enhancement and denoising method for multi-behavior recommendation
Xuan Zhang 0002, Weiyi Shang, Yubin Ma, Zhi Jin 0001 |
Neural Networks | 3 |
| 2025 | SQGE: Support-query prototype guidance and enhancement for few-shot relational triple extraction
Chen Gao 0006, Xuan Zhang 0002, Zhi Jin 0001, Kunpeng Du, Chunlin Yin, Tong Li 0004 |
Neural Networks | 2 |
| 2025 | Patient teacher can impart locality to improve lightweight vision transformer on small dataset
Jun Ling, Xuan Zhang 0002, LinYu Li 0001, Weiyi Shang, Chen Gao 0006, Tong Li 0004 |
Pattern Recognit. | 2 |
| 2025 | PRPSV: Parking Efficiency and Reservation Service Optimization Based on Parking Space ViewabstractDifficulty in parking leads to many issues, such as traffic congestion, and hinders the development of intelligent transportation systems. One significant reason is that the cruise parking does not fully utilize real-time information about the parking lot (e.g., availability status of the parking spaces), resulting in low parking efficiency and high parking costs. Even though reservation parking improves parking efficiency, its reservation service is coarse (e.g., could not reserve a specific parking space). Therefore, to achieve more efficient parking and optimize existing reservation services, we propose a real-time parking space view (PSV) framework (called PRPSV). PSV can reflect the position distribution and availability status of each parking space in a parking lot, which enables drivers to quickly and efficiently obtain real-time parking information and complete parking decisions. However, there have been fewer studies on PSV in recent years, and these methods are high-cost, limited scalability, and do not use PSV to optimize parking efficiency and services. Therefore, we propose a method to construct and update PSV accurately. Further, we model cruise and reservation modes in non-PSV and PSV-based scenarios to compare and analyze the impact of PSV on parking efficiency. Finally, the comprehensive qualitative comparison with related work demonstrates the innovativeness of PRPSV and the sufficient experimental results and a case study in an actual parking lot show that PRPSV can efficiently and accurately construct and update PSV, and the introduction of PSV can effectively improve parking efficiency and optimize reservation services. Jishu Wang, Xuan Zhang 0002, LinYu Li 0001, Xue Wang 0011, Shenglong Lv, Rui Zhu 0009, Tong Li 0004 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Multi-View Riemannian Manifolds Fusion Enhancement for Knowledge Graph CompletionabstractAs the application of knowledge graphs becomes increasingly widespread, the issue of knowledge graph incompleteness has garnered significant attention. As a classical type of non-euclidean spatial data, knowledge graphs possess various complex structural types. However, most current knowledge graph completion models are developed within a single space, which makes it challenging to capture the inherent knowledge information embedded in the entire knowledge graph. This limitation hinders the representation learning capability of the models. To address this issue, this paper focuses on how to better extend the representation learning from a single space to Riemannian manifolds, which are capable of representing more complex structures. We propose a new knowledge graph completion model called MRME-KGC, based on multi-view Riemannian Manifolds fusion to achieve this. Specifically, MRME-KGC simultaneously considers the fusion of four views: two hyperbolic Riemannian spaces with negative curvature, a Euclidean Riemannian space with zero curvature, and a spherical Riemannian space with positive curvature to enhance knowledge graph modeling. Additionally, this paper proposes a contrastive learning method for Riemannian spaces to mitigate the noise and representation issues arising from Multi-view Riemannian Manifolds Fusion. This paper presents extensive experiments on MRME-KGC across multiple datasets. The results consistently demonstrate that MRME-KGC significantly outperforms current state-of-the-art models, achieving highly competitive performance even with low-dimensional embeddings. LinYu Li 0001, Zhi Jin 0001, Xuan Zhang 0002, Haoran Duan 0002, Jishu Wang, Zhengwei Tao, Haiyan Zhao 0001, Xiaofeng Zhu 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | RLChain: A DRL Approach for Blockchain Performance Optimization Toward IIoTabstractWith the development of communication technology and Internet of Things, Industrial Internet of Things (IIoT) is proposed in the automation industry for complex scenarios. Blockchain is applied in IIoT to solve data security and privacy issues related to centralized data storage and processing. However, there are inevitably performance issues with throughput constraints when blockchain manages large amounts of device data. This paper proposes a blockchain-supported performance optimization framework for IIoT systems using deep reinforcement learning (DRL) methods. We model the blockchain performance optimization problem as a Markov decision process that optimizes the blockchain’s throughput by dynamically adjusting the block size and interval through DRL while satisfying security constraints. We use the double deep Q-network (DDQN) to deal with the dynamic and complexity of optimization problems due to the heterogeneity of equipment and diversified requirements. We also alleviate the overestimation problem caused by DQN. Meanwhile, we study the impact of the number of network layers and different activation units on the performance optimization method in DDQN. Finally, we prove that our work is feasible and effective through the case study based on actual IIoT scenario datasets. Experimental results demonstrate that our proposed scheme enhances blockchain performance in IIoT systems. The detailed qualitative comparison with related work demonstrates the superiority and innovation of our work and proves that it improves the shortcomings of existing work. Min An, Xuan Zhang 0002, Jishu Wang, Qiyuan Fan, Chen Gao 0006, LinYu Li 0001, Cuizhen Lu, Yingchen Liu |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2025 | Fed-OLF: Federated Oversampling Learning Framework for Imbalanced Software Defect Prediction Under Privacy ProtectionabstractSoftware defect prediction technology can discover potential errors or hidden defects by establishing prediction models before the use of products in the field of software engineering, so as to reduce subsequent problems and improve software quality and security. However, building predictive models requires enough software defect dataset support, especially defect samples. Due to the involvement of confidential information from various organizations or enterprises, software defect data cannot be shared and effectively utilized. Therefore, to achieve collaborative training of multiparty shared software defect prediction models while keeping the data local to various organizations, we made the federated learning framework for the issue of software defect prediction. Meanwhile, the nondefect and defect instances in software defect datasets are usually imbalanced, which can seriously affect the software defect prediction performance of the model. Therefore, this study designs a novel federated oversampling learning framework Fed-OLF. First, the TabDiT method based on deep generative model is proposed in Fed-OLF to expand and rebalance the local imbalanced software defect dataset of each client with a certain degree of privacy protection. Second, a parameter aggregation strategy based on local information entropy is proposed in Fed-OLF to further optimize the parameter aggregation effect of the global shared model, thereby achieving better model performance. We conduct extensive experiments on the PROMISE dataset and the NASA Promise repository, and experimental results on the PROMISE dataset and the NASA Promise repository show that, the proposed Fed-OLF exhibits better predictive performance under the F1-score, G-mean, and AUC metrics when compared with the advanced baseline methods. In addition, we verify that both the TabDiT method and the parameter aggregation strategy based on local information entropy in Fed-OLF are useful, and the combination of them can more effectively improve model performance. Ming Zheng, Rui Zhu 0009, Xuan Zhang 0002, Zhi Jin 0001 |
IEEE Trans. Reliab. | 4 |
| 2024 | Temporal knowledge graph reasoning based on evolutional representation and contrastive learning
Qiuying Ma, Xuan Zhang 0002, Zishuo Ding, Chen Gao 0006, Weiyi Shang, Qiong Nong, Yubin Ma, Zhi Jin 0001 |
Appl. Intell. | 2 |
| 2024 | An estimation method for multidimensional urban street walkability based on panoramic semantic segmentation and domain adaptation
Xuan Zhang 0002, LinYu Li 0001, Chen Gao 0006, Jun Ling |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | PowerPulse: Power energy chat model with LLaMA model fine-tuned on Chinese and power sector domain knowledgeabstractAbstract Recently, large‐scale language models (LLMs) such as chat generative pre‐trained transformer and generative pre‐trained transformer 4 have demonstrated remarkable performance in the general domain. However, inadaptability in a particular domain has led to hallucination for these LLMs when responding in specific domain contexts. The issue has attracted widespread attention, existing domain‐centered fine‐tuning efforts have predominantly focused on sectors like medical, financial, and legal, leaving critical areas such as power energy relatively unexplored. To bridge this gap, this paper introduces a novel power energy chat model called PowerPulse. Built upon the open and efficient foundation language models (LLaMA) architecture, PowerPulse is fine‐tuned specifically on Chinese Power Sector Domain Knowledge. This work marks the inaugural application of the LLaMA model in the field of power energy. By leveraging pertinent pre‐training data and instruction fine‐tuning datasets tailored for the power energy domain, the PowerPulse model showcases exceptional performance in tasks such as text generation, summary extraction, and topic classification. Experimental results validate the efficacy of the PowerPulse model, making significant contributions to the advancement of specialized language models in specific domains. Chunlin Yin, Kunpeng Du, Qiong Nong, Hongcheng Zhang, Xuan Zhang 0002 |
Expert Syst. J. Knowl. Eng. | 9 |
| 2024 | GIMM: A graph convolutional network-based paraphrase identification model to detecting duplicate questions in QA communities
Kunpeng Du, Xuan Zhang 0002, Chen Gao 0006, Rui Zhu 0009, Qiong Nong, XianYu Yang, Chunlin Yin |
Multim. Tools Appl. | 2 |
| 2024 | Fine-grained cybersecurity entity typing based on multimodal representation learning
Baolei Wang, Xuan Zhang 0002, Jishu Wang, Chen Gao 0006, Qing Duan, LinYu Li 0001 |
Multim. Tools Appl. | 2 |
| 2024 | Few-shot relational triple extraction with hierarchical prototype optimization
Chen Gao 0006, Xuan Zhang 0002, Zhi Jin 0001, Weiyi Shang, Yubing Ma, LinYu Li 0001, Zishuo Ding, Yuqin Liang |
Pattern Recognit. | 2 |
| 2024 | PEAE-GNN: Phishing Detection on Ethereum via Augmentation Ego-Graph Based on Graph Neural NetworkabstractRecent years, the successful application of blockchain in cryptocurrency has attracted a lot of attention, but it has also led to a rapid growth of illegal and criminal activities. Phishing scams have become the most serious type of crime in Ethereum. Some existing methods for phishing scams detection have limitations, such as high complexity, poor scalability, and high latency. In this article, we propose a novel framework named phishing detection on Ethereum via augmentation ego-graph based on graph neural network (PEAE-GNN). First, we obtain account labels and transaction records from authoritative websites and extract ego-graphs centered on labeled accounts. Then we propose a feature augmentation strategy based on structure features, transaction features and interaction intensity to augment the node features, so that these features of each ego-graph can be learned. Finally, we present a new graph-level representation, sorting the updated node features in descending order and then taking the mean value of the top n to obtain the graph representation, which can retain key information and reduce the introduction of noise. Extensive experimental results show that PEAE-GNN achieves the best performance on phishing detection tasks. At the same time, our framework has the advantages of lower complexity, better scalability, and higher efficiency, which detects phishing accounts at early stage. Xuan Zhang 0002, Jishu Wang, Chen Gao 0006, Rui Zhu 0009, Qiuying Ma |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2024 | LearningChain: A Highly Scalable and Applicable Learning-Based Blockchain Performance Optimization FrameworkabstractBlockchain is a trans-generational technology that is gradually introduced and applied in many fields because of its characteristics such as tamper-proof, traceability, and decentralization. However, the performance bottlenecks of blockchain have been one factor that hinders its practical application. This paper proposes a blockchain performance optimization framework (called LearningChain). We use a temporal convolution network to predict the transaction arrival rate of the blockchain and propose an ensemble learning-based method and a meta-learning-based method to train a blockchain performance prediction model, respectively. We design a performance scoring mechanism to dynamically tune the configuration parameters of the blockchain to optimize the blockchain performance. In addition, we collect and contribute a blockchain performance dataset (called HFBTP) for other researchers to research. The sufficient experimental results and analysis show that LearningChain can effectively optimize blockchain performance. The quantitative and qualitative comparisons with related work demonstrate the superiority and innovation of our work, LearningChain reaches state-of-the-art, is highly applicable, scalable, and can be applied to many practical blockchain-based application scenarios and different blockchain platforms. LearningChain can be complemented with other existing blockchain performance optimization tools and methods to further enhance the effectiveness of blockchain performance optimization. Jishu Wang, Xuan Zhang 0002, Zhi Jin 0001, LinYu Li 0001, Rui Zhu 0009, Shenglong Lv |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2024 | Business Process Retrieval From Large Model Repositories for Industry 4.0abstractThe process model repository has demonstrated unprecedented success in a variety of industrial and process as a service scenarios. With the rapid increase of massive business process-related data under Industry 4.0, effectively retrieval of process models from large process model repositories becomes a critical challenge for process mining, process deployment and process model acquisition. To accelerate the retrieval of process models from a large process repository, existing retrieval methods rely solely on building single dimension process model indices. In this article we show that this single dimension indexing approach is not only inefficient but also cumbersome for supporting high performance retrieval services over large process model repositories. We propose a new business process model indexing and retrieval with structure and behavior fusion. In the indexing stage, we propose a process model index generation paradigm method with two novel features. First, our index algorithm can transform thetrace equivalent process model(TEPM) with complex structures into a process tree, which can better capture process sequence semantics than the existing approach based on block structured process model. Second, we improve the method for computing the process tree edit distance for measuring process model similarity by introducing the process tree similarity method, which can distinguish leaf nodes and non-leaf nodes and improve the limitations of the traditional edit distance algorithm. Extensive experiments using real world process repositories demonstrate that the proposed methods are under polynomial time in both the model index generation and model querying stages, and offer superior retrieval performance compared to existing process model retrieval methods in terms of efficiency, search capability and scope. Rui Zhu 0009, Ling Liu 0001, Wei Zhou 0011, Xuan Zhang 0002, Yeting Chen |
IEEE Trans. Serv. Comput. | 5 |
| 2023 | Multimodal Sentiment Analysis under modality deficiency with prototype-Augmentation in software engineeringabstractSentiment analysis has a wide range of promising applications in software engineering, and the development of deep learning has demonstrated that the uniform representation of different modalities can improve the model performance of sentiment analysis. However, in practical applications, multimodal sentiment analysis always faces unsatisfactory situations, especially when the modality has missing samples, most models may fail. For example, social dynamics of technicians in developer communities can face modality unavailability due to privacy settings. Several existing works based on deep learning and regularization methods have explored the modal missing problem, but these works cannot balance the cases of modal general missing (rate < 50%) and severe missing (rate ≥ 50%), and do not consider the resource consumption during model inference. Therefore, in this paper, we proposed a prototype augmented multimodal teacher-student network (PAMD) to address the above issues. Specifically, a multi-level and multi-origin distillation strategy is used to minimize the required resources and inference time, and prototype augmentation is used to guarantee the performance of the model when a modality is severely missing. Extensive experiments are conducted on different benchmark datasets to explore a network that balances performance and resource consumption. And It achieves good results in different modalities of missing cases. Baolei Wang, Xuan Zhang 0002, Kunpeng Du, Chen Gao 0006, LinYu Li 0001 |
SANER | 2 |
| 2023 | Enhancing recommendations with contrastive learning from collaborative knowledge graph
Yubin Ma, Xuan Zhang 0002, Chen Gao 0006, Yahui Tang, LinYu Li 0001, Rui Zhu 0009, Chunlin Yin |
Neurocomputing | 2 |
| 2023 | Knowledge graph completion method based on quantum embedding and quaternion interaction enhancement
LinYu Li 0001, Xuan Zhang 0002, Zhi Jin 0001, Chen Gao 0006, Rui Zhu 0009, Yuqin Liang, Yubing Ma |
Inf. Sci. | 2 |
| 2023 | ERGM: A multi-stage joint entity and relation extraction with global entity match
Chen Gao 0006, Xuan Zhang 0002, LinYu Li 0001, JinHong Li, Rui Zhu 0009, Kunpeng Du, Qiuying Ma |
Knowl. Based Syst. | 2 |
| 2023 | BPR: Blockchain-Enabled Efficient and Secure Parking Reservation Framework With Block Size Dynamic Adjustment MethodabstractThe parking lot is one of the important components of the intelligent transportation system (ITS). The current parking lots mainly use instant parking, which has low parking efficiency, during peak hours, which leads to traffic congestion. To guarantee the stable operation of parking lots, we propose a blockchain-enabled parking reservation framework, called BPR. Traditional parking reservation systems may exist the condition of malicious reservations, and resulting in wasted parking spaces. Therefore, we design a reputation mechanism to manage the parking reservation behavior of vehicles and reduce the number of malicious nodes. In addition, to balance the performance of the blockchain at different times (especially during peak hours), we use deep learning (DL) to dynamically adjust the block size to make the blockchain run more efficiently and stably. We deploy the system in Hyperledger Fabric and conduct effectiveness experiments. The comprehensive evaluation results and analysis show that the proposed reputation mechanism can effectively curb malicious nodes from reserving parking spaces and reduce the waste of parking resources. And the block size will be dynamically adjusted to balance the performance of the blockchain at different periods, this method is also applicable to other blockchain performance-sensitive scenes. Finally, this paper is compared with related work to demonstrate the innovation and feasibility of this work from various aspects. Jishu Wang, Chen Miao, Rui Zhu 0009, Xuan Zhang 0002, Yahui Tang, Chen Gao 0006 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | KG2Lib: knowledge-graph-based convolutional network for third-party library recommendation
Zhao Jingzhuan, Xuan Zhang 0002, Chen Gao 0006, Zhudong Li, Bao-lei Wang |
J. Supercomput. | 2 |
| 2022 | A knowledge graph completion model based on contrastive learning and relation enhancement method
LinYu Li 0001, Xuan Zhang 0002, Yubin Ma, Chen Gao 0006, Jishu Wang, Yong Yu 0009, Qiuying Ma |
Knowl. Based Syst. | 2 |
| 2021 | A survey on the techniques, applications, and performance of short text semantic similarityabstractSummary Short text similarity plays an important role in natural language processing (NLP). It has been applied in many fields. Due to the lack of sufficient context in the short text, it is difficult to measure the similarity. The use of semantics similarity to calculate textual similarity has attracted the attention of academia and industry and achieved better results. In this survey, we have conducted a comprehensive and systematic analysis of semantic similarity. We first propose three categories of semantic similarity: corpus‐based, knowledge‐based, and deep learning (DL)‐based. We analyze the pros and cons of representative and novel algorithms in each category. Our analysis also includes the applications of these similarity measurement methods in other areas of NLP. We then evaluate state‐of‐the‐art DL methods on four common datasets, which proved that DL‐based can better solve the challenges of the short text similarity, such as sparsity and complexity. Especially, bidirectional encoder representations from transformer model can fully employ scarce information of short texts and semantic information and obtain higher accuracy and F1 value. We finally put forward some future directions. Mengting Han, Xuan Zhang 0002, Wei Yun, Chen Gao 0006 |
Concurr. Comput. Pract. Exp. | 2 |
| 2021 | Data and knowledge-driven named entity recognition for cyber securityabstractAbstract Named Entity Recognition (NER) for cyber security aims to identify and classify cyber security terms from a large number of heterogeneous multisource cyber security texts. In the field of machine learning, deep neural networks automatically learn text features from a large number of datasets, but this data-driven method usually lacks the ability to deal with rare entities. Gasmi et al. proposed a deep learning method for named entity recognition in the field of cyber security, and achieved good results, reaching an F1 value of 82.8%. But it is difficult to accurately identify rare entities and complex words in the text.To cope with this challenge, this paper proposes a new model that combines data-driven deep learning methods with knowledge-driven dictionary methods to build dictionary features to assist in rare entity recognition. In addition, based on the data-driven deep learning model, an attention mechanism is adopted to enrich the local features of the text, better models the context, and improves the recognition effect of complex entities. Experimental results show that our method is better than the baseline model. Our model is more effective in identifying cyber security entities. The Precision, Recall and F1 value reached 90.19%, 86.60% and 88.36% respectively. Chen Gao 0006, Xuan Zhang 0002, Hui Liu 0061 |
Cybersecur. | 2 |
| 2021 | Knowledge modeling: A survey of processes and techniquesabstractKnowledge modeling is an important step in building knowledge-based applications. Understanding the processes of knowledge modeling and the techniques involved can help developers to grasp the knowledge modeling task as a whole and improve the efficiency of execution and management of modeling tasks. However, previous reviews on knowledge modeling mainly focus on ontology-based knowledge modeling. At present, there is no research work to summarize nonontology knowledge modeling methods, nor to systematically summarize the processes and techniques of knowledge modeling. In this paper, the processes, techniques, and characteristics of knowledge modeling methods based on ontology and nonontology are surveyed. Three research questions related to knowledge modeling are proposed. (1) What methods can be used for knowledge modeling? (2) What processes are involved in knowledge modeling? (3) What techniques are used in the processes of knowledge modeling? By answering these questions, the results of the survey help developers choose appropriate knowledge modeling methods in their work and complete modeling tasks effectively. Meanwhile, it is also conducive to the research work of improving knowledge modeling methods in the future. Wei Yun, Xuan Zhang 0002, Zhudong Li, Hui Liu 0061, Mengting Han |
Int. J. Intell. Syst. | 2 |
| 2021 | A review on cyber security named entity recognitionabstractWith the rapid development of Internet technology and the advent of the era of big data, more and more cyber security texts are provided on the Internet. These texts include not only security concepts, incidents, tools, guidelines, and policies, but also risk management approaches, best practices, assurances, technologies, and more. Through the integration of large-scale, heterogeneous, unstructured cyber security information, the identification and classification of cyber security entities can help handle cyber security issues. Due to the complexity and diversity of texts in the cyber security domain, it is difficult to identify security entities in the cyber security domain using the traditional named entity recognition (NER) methods. This paper describes various approaches and techniques for NER in this domain, including the rule-based approach, dictionary-based approach, and machine learning based approach, and discusses the problems faced by NER research in this domain, such as conjunction and disjunction, non-standardized naming convention, abbreviation, and massive nesting. Three future directions of NER in cyber security are proposed: (1) application of unsupervised or semi-supervised technology; (2) development of a more comprehensive cyber security ontology; (3) development of a more comprehensive deep learning model. Chen Gao 0006, Xuan Zhang 0002, Mengting Han, Hui Liu 0061 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2020 | Pattern-based software process modeling for dependabilityabstractAbstract Traditional process modeling focuses on modeling activities for functional requirements. For dependability requirements, a knowledge‐based aspect‐oriented software process modeling approach is proposed. First, we extend the pattern and context to the knowledge graph triplet structure to describe dependability‐oriented knowledge patterns. By applying the patterns, dependability requirements can be organized into dependability‐related activities that are integrated into the software process. Then, aspect‐oriented modeling patterns based on Petri nets are introduced to support the integration of these dependability‐related activities and model dependability‐oriented software processes. Finally, the modeling performance and the subjective usability of the patterns are evaluated by 110 students with different degrees. The results indicate that these two indexes are on the positive track. Hence, the patterns may be the backbone of dependability‐oriented software process modeling. Xuan Zhang 0002, Wei Yun, Chen Gao 0006, Mengting Han, Hui Liu 0061 |
J. Softw. Evol. Process. | 1 |
| 2018 | Trustworthiness requirement-oriented software process modelingabstractAbstract Trustworthy software is delivered by enacting trustworthy software processes. The purpose of this paper is to propose an approach to modeling trustworthiness requirement‐oriented software processes. First, based on the aspect‐oriented modeling techniques, separation of concerns is used to separate the crosscutting activities and the core activities according to the different trustworthiness requirements and functional requirements. A goal‐oriented modeling and reasoning method for trustworthiness requirements to find the crosscutting activities that satisfy multiple trustworthiness requirements is presented. Then, base processes are modeled for functional requirements. The crosscutting activities for trustworthiness requirements are decomposed into processes or tasks and encapsulated in aspects that are woven into the base processes. In the weaving procedure, correct weaving methods between multiple aspects and between aspects and base processes are designed. Errors or mistakes of aspect‐oriented process modeling are prevented. Finally, trustworthy third‐party certification authority software is studied systematically in a case study, and performance evaluations are conducted to show the cost and effect of the approach. Xuan Zhang 0002, Yanni Kang |
J. Softw. Evol. Process. | 1 |
| 2013 | Completeness set proof of precondition and post-condition types of activity in any EPMabstractSoftware evolution process model (EPM) is created in terms of a formal evolution process meta-model (EPMM) and semi-formal approach to modeling based on EPMM [1]. In order to better manage and control the software evolution process and make the best of existing software technology, the method to transform any EPM to its execution model based logic programming has been proposed. Completeness of conversion depends on completeness of the rules, that is, all the expressions of the original model are found the correspondence in the target model. Since transformation rules are proposed based on precondition or post-condition types of activities in anyone EPM, this need to prove that activity type set in anyone EPM is completeness set. To this end, the precondition and post-condition of activities in EPM are classified based on analyzing all expressions in EPMs and the semantics of the activity execution. Type completeness set of activity’s precondition and its post-condition is presented. Lastly we prove that the activity type set in anyone EPM is completeness set by mathematical induction. Tong Li 0004, Jinzhuo Liu, Xuan Zhang 0002, Yong Yu 0009 |
ICMV | 4 |