EDBT 2026 Demo / reviewers in the wild / expert
Yude Bai
dblp:254/6031
· DBLP profile ↗
33ranked-venue papers
3as first author
30since 2021 · last 2026
0000-0002-9768-8045ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 11 since 2021Software engineering, systems software and programming languages · 10 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Attention Network with Correction for Cross-Domain User AssociationabstractDespite the rich spatiotemporal patterns contained in trajectory data from multiple Location-Based Social Network (LBSN) platforms, heterogeneous formats, semantic inconsistencies, and unequal user scales across platforms create substantial barriers to reliable identity mapping. Furthermore, GPS drift and sparse sampling result in degraded data quality and distribution imbalance, which render existing trajectory representation methods inadequate for capturing high-order dependencies and dynamic spatiotemporal evolution patterns in heterogeneous multi-relational graphs. To this end, we propose HANCUA (Hierarchical Attention Network with Correction for User Association), a novel framework that employs a dual-stage correction mechanism to enhance cross-domain trajectory analysis. The approach constructs hierarchical multi-relational graphs comprising location, trajectory, and correction layers to capture fine-grained mobility patterns, behavioral associations, and inter-platform distribution differences. We design relation-aware multi-head graph attention networks to model complex interactions among heterogeneous node types, which enables comprehensive spatial relationship modeling. A spatiotemporal semantic collaborative learning module integrates temporal information with mobility patterns through interaction-aware attention mechanisms, while an ensemble correction decision module incorporates ensemble learning principles to systematically correct user association biases and address distribution imbalance problems. Extensive experiments on two real-world LBSN cross-domain datasets reveals that HANCUA significantly outperforms state-of-the-art methods in user identity linking accuracy. Chenlong Wu, Yude Bai |
AAAI | 4 |
| 2026 | BioLemons: Latent Conditional Diffusion Model with VAE Embedding for Enhancing Spatial Transcriptomics
Haolu Zhou, Wenying He, Yude Bai, Fei Guo 0001 |
DASFAA (3) | 3 |
| 2026 | Spatial-Semantic Attacks and Protection for Trajectory Privacy: From Vulnerability to Defense
Zhuo Han, Ze Wang 0016, Yude Bai, Ji Zhang 0001 |
ICC | 3 |
| 2026 | DGT-LLM: A Multimodal Industrial Signal Learning Framework for Petroleum Production Forecasting
Jianzhuo Liu, Yude Bai, Pengji Qian, Jin Hao |
ICIC (15) | 2 |
| 2026 | MUPSI: A Multi-party Updatable Delegated Private Set Intersection Protocol for Privacy-Preserving Medical Data Collaboration
Weihao Mo, Yude Bai, Yongyang Lv, Luxue Yu, Can Gao, Zimo Wang |
ICIC (15) | 2 |
| 2026 | POF-HG: Fusion of Public Opinion Field Effect and Heterogeneous Hypergraph for Information Diffusion Prediction
Yude Bai |
PAKDD (3) | 4 |
| 2026 | DualDis: A Dual Disentanglement Network for Vehicle Re-identification
Wenying He, Guangquan Xu, Yude Bai, Fei Guo 0001 |
WWW | 4 |
| 2026 | Rubato: Efficient Post-Quantum Asynchronous Distributed Randomness Beacon With Integrated ConsensusabstractDistributed randomness beacons are essential for distributed systems (e.g., blockchain and MPC), providing un biased and unpredictable shared randomness. However, implementations in asynchronous networks often suffer from poor scal ability, low throughput, and high resource consumption when deployed as independent protocols. The state-of-the-art HashRand (CCS'24) achieves high throughput and low computational in tensity using only lightweight post-quantum cryptographic primitives for an independent asynchronous beacon. Building further on this, we propose Rubato, a low-overhead, high-throughput beacon protocol that leverages lightweight batched Asynchronous Complete Secret Sharing with Byzantine Atomic Broadcast-based state machine replication (via our tailored RubatoSMR). Rubato reduces communication complexity by an O(clogn) factor compared to HashRand and resolves the circular dependency between beacon and BAB-SMR in asynchronous settings. Experiments on AWSdemonstrate that Rubato achieves ideal overall performance in scalability, resource usage, and throughput; for instance, at n = 121nodes, it produces an average of 174 beacons per minute, with RubatoSMR further optimizing memory and bandwidth consumption. Linghe Yang, Tonghong Chong, Jian Liu 0004, Jingyi Cui, Guangquan Xu, Yude Bai, Lei Zhang 0024, Tao Luo 0010 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2026 | Knowledge is Power: A Knowledge Graph-Based Approach for Mobile Malware Traceability Analysis
Yao Zhang 0019, Guangquan Xu, Xiaohong Li 0001, Sen Chen 0001, Zhenchang Xing, Yude Bai, Yongqiang Lyu 0001, Wei Gong 0001, Xibin Zhao |
IEEE Trans. Mob. Comput. | 7 |
| 2026 | LIMR: Intent-Aware Mashup API Recommendation via LLM-Augmented Multi-Scale FusionabstractThe increasing availability of Web APIs has amplified the complexity of mashup creation, where developers must identify compatible and functionally relevant APIs based on often ambiguous natural language descriptions. Traditional methods also fall short in capturing hierarchical semantic cues, modeling compatibility, and aligning with developer intent. Although large language models (LLMs) offer strong generalization capabilities, they remain unreliable in mashup recommendation due to hallucinated outputs, limited controllability, and token-length constraints when dealing with large-scale API repositories. To overcome these limitations, we introduceLIMR, an intent-aware mashup recommendation framework that combines LLM-augmented semantic reasoning with structured, multi-scale neural modeling.LIMRfirst prompts a LLM to extract high-level intent from user requirements, which serves as a global semantic signal. This intent is fused with low-level, multi-scale features extracted by a convolutional encoder, which are designed to capture fine-grained lexical/phrasal patterns at different granularities and provide precise semantic grounding for API matching. These heterogeneous representations are further contextually refined through a Transformer-based interaction module. To handle nonlinear semantic dependencies and compositional complexity,LIMRintegrates a Kolmogorov-Arnold Network (KAN) with learnable activation functions, enhancing the model's capacity to capture intricate feature interactions. The entire framework is optimized via LLM, incorporating auxiliary objectives such as mashup category prediction and API quality estimation to guide generalization and reduce overfitting. Comprehensive experiments on the ProgrammableWeb and APIBench datasets show thatLIMRsignificantly outperforms state-of-the-art baselines, which the ranking-oriented metrics, including NDCG and mAP, achieves improvements of 17.1%–34.2% over the strongest competitors. These results confirm the effectiveness ofLIMR's hybrid design in delivering precise, robust, and intent-aware mashup API recommendations, especially in scenarios where LLMs alone fail to meet accuracy and scalability demands. Yao Zhang 0019, Yude Bai, Minhong Dong, Keqing Cen, Ji Zhang 0001, Wei Ma 0014, Yongqiang Lyu 0001, Xiaohong Li 0001, Junjie Wang 0007, Lingxiao Jiang, Yang Liu 0003 |
IEEE Trans. Serv. Comput. | 2 |
| 2025 | Cross-Domain Trajectory Association Based on Hierarchical Spatiotemporal Enhanced Attention HypergraphabstractIdentifying and linking the same users across different social platforms is crucial for understanding user behavior and preferences. However, cross-domain datasets exhibit diverse characteristics, such as varying check-in frequencies, significant disparities in data precision, and distinct distributions. Existing trajectory representations rely on recurrent neural network, which fails to dynamically learn multi-dimensional feature relations and capture high-order associations. Furthermore, current methods for integrating trajectory information fails to capture the complex relations and dynamic variations among cross-domain mobility trajectories. To this end, we propose the Hierarchical Spatio-Temporal Enhanced Attention Hypergraph Network (StarNet). This model dynamically regulates the multi-dimensional features of trajectories through a locally enhanced spatiotemporal graph neural network. Meanwhile, StarNet employs a hypergraph network enhanced by a global spatiotemporal to capture high-order associations between cross-domain trajectories. The fusion enhancement association integrates local and global information, which enables this model to link user identities. Extensive experiments on two well-known LBSN cross-domain datasets reveal that StarNet outperforms state-of-the-art baselines in the accuracy of user identity linkage. Chenlong Wu, Keqing Cen, Yude Bai, Jin Hao |
AAAI | 4 |
| 2025 | LGATFormer: A Dual-Path Model Combining Line Graph Attention and Transformer for Gene Regulatory Network InferenceabstractReconstructing high-precision gene regulatory networks (GRNs) from single-cell RNA sequencing (scRNA-seq) data presents significant challenges, including high noise levels, data sparsity, and structural complexity. We introduce LGATFormer, a novel deep learning model that combines local structural modeling with global dependency extraction to address these challenges effectively. LGATFormer extracts enclosing subgraphs centered on target gene pairs, transforms them into line graphs, and uses Graph Attention Networks (GATs) to learn edge-level representations, capturing high-order regulatory structures. For global modeling, the model employs a Transformer encoder to process the entire gene expression matrix, utilizing self-attention mechanisms to model long-range dependencies between genes. Across multiple benchmarks, LGATFormer outperforms existing methods by achieving state-of-the-art AUROC and AUPRC on 89.29% of STRING and Non-Specific ground-truth networks, while exhibiting superior generalization, robustness, and interpretability. This model offers an effective and reliable solution for GRN inference, advancing both theoretical and practical applications in systems biology. Wenying He, Yaowei Zhu, Rentao Zhang, Haolu Zhou, Yude Bai, Fei Guo 0001 |
BIBM | 5 |
| 2025 | Catching mRNA's Hiddens Marks: A Dual-Path Network by Contrastive Learning for N4-acetylcytidine PredictionabstractN4-acetylcytidine (ac4C) is a crucial RNA modification associated with mRNA stability and translational efficiency. Accurate identification of ac4C sites is essential for understanding their regulatory functions. However, experimental detection remains expensive and labor-intensive. At the same time, existing computational models suffer from limited generalization and insufficient feature discrimination, especially in distinguishing subtle nucleotide patterns. In this work, we propose a deep learning model named SNN-ac4C, which is based on a contrastive learning-based neural network. The model integrates a dual-path structure that combines BiLSTM and Multi-Head Self-Attention (MHSA) for capturing long-range dependencies and global context, while using CNN to extract local biological sequence features. The contrastive learning module further enhances the discriminative ability of ac4C and Non-ac4C sites by increasing the separation between positive and negative samples. Experiments on the test set confirm the effectiveness of SNN-ac4C, which achieves an accuracy (ACC) of 84.60% and a Matthews Correlation Coefficient (MCC) of 0.6934. Compared with NBCR-ac4C, the current state-of-the-art model, SNN-ac4C improves ACC and MCC by 1.09% and 0.0228, respectively. The source code and relevant supplementary are publicly available at https://github.com/2103374200/SNN. Wenying He, Haolu Zhou, Yun Zuo 0001, Yude Bai, Fei Guo 0001 |
ECAI | 5 |
| 2025 | EfficientNet-BSFT-S: Dynamic Multi-Scale Modules for Robust Image Classification
Chengjie Guo, Minghong Dong, Xuewei Liu, Meng Xing, Yao Zhang 0019, Yude Bai |
ICIC (19) | 6 |
| 2025 | DCFICSH: A Dual-Channel Fusion Model Combining Multi-Modal Data for Identifying Cell-Specific Silencers and Their Strength in the Human Genome
Jingdong Yuan, Qinqin Zhu, Haolu Zhou, Yun Zuo 0001, Yude Bai, Wenying He |
ICIC (19) | 6 |
| 2025 | PneumoNeXt: A Multi-Scale Attention and Contrastive Learning Approach for Pneumonia Diagnosis
Lirong Zhang, Meng Xing, Yao Zhang 0019, Yude Bai |
ICIC (19) | 4 |
| 2025 | Filling the Missings: Spatiotemporal Data Imputation by Conditional DiffusionabstractMissing data in spatiotemporal systems presents a significant challenge for modern applications, ranging from environmental monitoring to urban traffic management. The integrity of spatiotemporal data often deteriorates due to hardware malfunctions and software failures in real-world deployments. Current approaches based on machine learning and deep learning struggle to model the intricate interdependencies between spatial and temporal dimensions effectively and, more importantly, suffer from cumulative errors during the data imputation process, which propagate and amplify through iterations. To address these limitations, we propose CoFILL, a novel Conditional Diffusion Model for spatiotemporal data imputation. CoFILL builds on the inherent advantages of diffusion models to generate high-quality imputations without relying on potentially error-prone prior estimates. It incorporates an innovative dual-stream architecture that processes temporal and frequency domain features in parallel. By fusing these complementary features, CoFILL captures both rapid fluctuations and underlying patterns in the data, which enables more robust imputation. The extensive experiments demonstrate that CoFILL's noise prediction network successfully transforms random noise into meaningful values that align with the true data distribution. The results also show that CoFILL outperforms state-of-the-art methods in terms of imputation accuracy. The source code is publicly available at https://github.com/joyHJL/CoFILL. Wenying He, Jieling Huang, Junhua Gu, Ji Zhang 0001, Yude Bai |
IJCAI | 5 |
| 2025 | Long- and Short-Term Feature Fusion Network for River Velocity Forecasting: From Past to FutureabstractAccurate river velocity prediction plays a vital role in water resource management and hydraulic engineering. Seasonal variations and weather conditions introduce complex temporal patterns, making prediction a challenging task. Traditional methods struggle to account for both short-term fluctuations and long-term trends, and they fail to fully exploit temporal information. To overcome these limitations, we propose SeekRiver, a prediction model that integrates both long- and short-term features. We first transform time features into low-dimensional vectors through an embedding layer, enhancing their temporal representation. The model uses two LSTM layers to separately capture short-term and long-term dependencies. An attention mechanism is employed to emphasize crucial features while minimizing the impact of irrelevant ones. A weighted fusion strategy then combines multi-scale features to improve prediction accuracy. Extensive experiments, including comparisons with baseline models and ablation studies, validate SeekRiver’s effectiveness. The model significantly outperforms traditional methods, with the attention mechanism and short-term feature module contributing the most to the performance improvement, which offers a robust and efficient solution for river velocity prediction. Ruijie Gong, Jinping Xie, Kaitao Liu, Minhong Dong, Yude Bai |
IJCNN | 7 |
| 2025 | Calmdroid: Core-Set Based Active Learning for Multi-Label Android Malware DetectionabstractOne of the trends in the evolution of Android malware is the increasing diversity of malicious behaviors, such as SMSrelated and Internet-related actions. Traditional binary or familybased classification methods are inadequate for fine-grained detection of these behaviors. Thus, multi-label classification is required to identify various malicious behaviors within a single malware sample. This paper employs an active learning strategy to add multi-behavior labels to large-scale datasets based on expert-annotated small-scale datasets. To address the issue of noisy labels (simulating real-world mislabeling), we propose CalmDroid, an active learning framework utilizing the coreset strategy, instead of the confuse-set strategy for updating the model with out-of-distribution (OOD) points. We evaluate CalmDroid's performance using the Drebin and VirusShare datasets. Experimental results demonstrate that CalmDroid achieves superior detection performance under varying noise conditions, with an accuracy improvement of up to 0.704 compared to the confuse-set strategy. In high-noise environments (15%), it reaches detection accuracy as high as 0.944. Additionally, we validate CalmDroid's capability to detect evolving malware. Despite behavioral evolution in Drebin malware across different time steps, CalmDroid consistently achieves detection rates above 70 % in the newest time step. Minhong Dong, Wenying He, Ze Wang 0016, Yude Bai |
ICPC | 7 |
| 2025 | Dummy-Trajectory Synthesis: A Privacy-Preserving Approach for Semantic Trajectory Data in IoT-Based LBSNabstractTrajectory data analysis is crucial in various applications but presents significant privacy risks, as location data can reveal sensitive information. Existing privacy protection methods, such as spatiotemporal K-anonymity and L-diversity, are vulnerable to semantic inference attacks, where public data is exploited to re-identify users. To address these challenges, we propose dummy-trajectory synthesis (DTS), an efficient privacy protection scheme for location-based social networks (LBSNs). DTS enhances privacy by leveraging users’ frequent behavioral sequences to generate synthetic dummy trajectories. Unlike traditional approaches, DTS considers both geographic and semantic data by segmenting historical trajectories into time periods using the OPTICS clustering algorithm. This enables the identification of regions with specific semantic attributes and the mining of semantic trajectory sequences. DTS optimizes dummy trajectory generation by combining Euclidean distances and semantic similarity, ranking historical points and establishing transition relationships. Experimental results show that DTS significantly improves privacy protection and performance compared to existing methods, without compromising service quality. DTS offers a robust solution for protecting trajectory data privacy in LBSNs against transition probability attack for joint time periods. Minhong Dong, Ze Wang 0016, Zhuo Han, Yude Bai, Xiaohu Ye, Guangquan Xu, Naixue Xiong |
IEEE Internet Things J. | 5 |
| 2025 | PEFN: A Patches Enhancement and Hierarchical Fusion Network for Robust Vehicle ReidentificationabstractVehicle Re-Identification (Re-ID), which is a significant application in the Internet of Things, aims to accurately retrieve the remaining images of a given vehicle across different cameras views. The improvement in vehicle Re-ID performance largely stems from better addressing the issues of inter-class similarity and intra-class variance. Existing methods, relying solely on max or average pooling after using attention modules, fail to obtain significantly complete and pure global and local features, and neglect the false guidance that some unique individual information on images bring to re-identification. Moreover, models combining global and local features have shown good results in vehicle Re-ID, but these successes neglect the interaction between features across different convolutional layers, resulting in the loss of crucial details for vehicle Re-ID. To tackle these issues, we introduce a Patches Enhancement and hierarchical Fusion Network (PEFN) based on a multi-branch architecture, divided into a Global and Local Attention Supplement (GLAS) branch, and an Enhanced Hierarchical feature fusion (EnHi) branch. The GLAS branch, through the Identity-related Feature Remodeling (IDFR) module’s staged supplementation of spatial and channel features, has achieved the enhancement of both global and local features and effectively mitigated the negative impacts of individual information. The EnHi branch enhances the robustness of feature representation by interacting hierarchical features. Extensive experiments on two large-scale vehicle re-identification datasets demonstrate that our PEFN method outperforms state-of-the-art vehicle re-identification approaches. Specifically, without utilizing extra data and re-ranking, our model achieves 85.15% mAP on the VeRi776 dataset. Code is available at https://github.com/711L/PEFN. Wenying He, Yude Bai, Naixue Xiong, Guangquan Xu, Fei Guo 0001 |
IEEE Internet Things J. | 3 |
| 2024 | Emotional Polarity Attention Mechanism for Text Sentiment Analysis
Tianheng Wang, Yude Bai |
DASFAA (5) | 4 |
| 2024 | TGSA: Trajectory Group Semantic Anonymization
Minhong Dong, Ze Wang 0016, Zhuo Han, Yude Bai, Guoying Qiu |
SecureComm (4) | 4 |
| 2023 | ArgusDroid: detecting Android malware variants by mining permission-API knowledge graph
Yude Bai, Sen Chen 0001, Zhenchang Xing, Xiaohong Li 0001 |
Sci. China Inf. Sci. | 1 |
| 2022 | Can Deep Learning Models Learn the Vulnerable Patterns for Vulnerability Detection?abstractDeep learning has been widely used for the security issue of vulnerability prediction. However, it is confusing to explain how a deep learning model makes decisions on the prediction, although such a model achieves a good performance. Meanwhile, it is also difficult to discover which part of the source code is concentrated on by this black-box model. To this end, we present an empirical evaluation to explore how the deep learning model works on predicting vulnerability and whether it precisely captures the critical code segments to represent the vulnerable patterns. First of all, we build a new vulnerability dataset, called Juliet+, in which vulnerability-related code lines of both positive (bad) and negative (good) samples are labeled manually with substantial efforts, based on the Juliet Test Suite. After that, four deep learning models by leveraging attention mechanisms are empirically implemented to detect vulnerability through mining vulnerable patterns from the source code. We conduct extensive experiments to evaluate the effectiveness of such four models and to analyze the interpretability with evaluation metrics such as Hit@k. The empirical experiment results reveal that the deep learning models with attention, to some extent, can focus on the vulnerability-related code segments that are profitable to interpret the result of vulnerability detection, especially when we adopt the graph neural network model. We further investigate what factors affect the interpretability of models including the class distribution, the number of samples, and the differences of sample features. We find the graph neural network model performs better on part of the dataset which contains balanced and sufficient samples with obvious differences between vulnerable and non-vulnerable patterns. Guoqing Yan, Sen Chen 0001, Yude Bai, Xiaohong Li 0001 |
COMPSAC | 3 |
| 2022 | Detecting and Augmenting Missing Key Aspects in Vulnerability DescriptionsabstractSecurity vulnerabilities have been continually disclosed and documented. For the effective understanding, management, and mitigation of the fast-growing number of vulnerabilities, an important practice in documenting vulnerabilities is to describe the key vulnerability aspects, such as vulnerability type, root cause, affected product, impact, attacker type, and attack vector. In this article, we first investigate 133,639 vulnerability reports in the Common Vulnerabilities and Exposures (CVE) database over the past 20 years. We find that 56%, 85%, 38%, and 28% of CVEs miss vulnerability type, root cause, attack vector, and attacker type, respectively. By comparing the differences of the latest updated CVE reports across different databases, we observe that 1,476 missing key aspects in 1,320 CVE descriptions were augmented manually in the National Vulnerability Database (NVD) , which indicates that the vulnerability database maintainers try to complete the vulnerability descriptions in practice to mitigate such a problem. To help complete the missing information of key vulnerability aspects and reduce human efforts, we propose a neural-network-based approach called PMA to predict the missing key aspects of a vulnerability based on its known aspects. We systematically explore the design space of the neural network models and empirically identify the most effective model design in the scenario. Our ablation study reveals the prominent correlations among vulnerability aspects when predicting. Trained with historical CVEs, our model achieves 88%, 71%, 61%, and 81% in F1 for predicting the missing vulnerability type, root cause, attacker type, and attack vector of 8,623 “future” CVEs across 3 years, respectively. Furthermore, we validate the predicting performance of key aspect augmentation of CVEs based on the manually augmented CVE data collected from NVD, which confirms the practicality of our approach. We finally highlight that PMA has the ability to reduce human efforts by recommending and augmenting missing key aspects for vulnerability databases, and to facilitate other research works such as severity level prediction of CVEs based on the vulnerability descriptions. Sen Chen 0001, Zhenchang Xing, Xiaohong Li 0001, Yude Bai, Jiamou Sun |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2021 | Key Aspects Augmentation of Vulnerability Description based on Multiple Security DatabasesabstractCommon Vulnerabilities and Exposures (CVE) is one of the most influential security databases. With the continuous disclosure of security vulnerabilities, the characteristics of them are documented as vulnerability reports, and it is discovered that their impact on computer systems is increasing. However, as our research continues to deepen, we find that the current lack of key aspects of CVE description is more serious than before. In response to this situation, our research focuses on how to correctly and completely extract key aspect descriptions from various security vulnerability databases to supplement the CVE reports. First, we fetch almost all of semi-structured vulnerability reports from the CVE, Security Focus, and IBM X-Force Exchange databases before November 2020. We then propose a customized NER (Named entity recognition) method based on deep neural networks to extract six key aspects from unstructured descriptions. Finally, we use the corresponding security vulnerability reports in other vulnerability databases to complete the missing key aspects in the correlated CVE description. We conduct sampling surveys on various aspects of this information, and verify the accuracy of extracting key aspects, and find that our method can extract key information from vulnerability descriptions. To demonstrate the usefulness of key aspects augmentation, after completing the missing affected product, root cause, attacker type, attack vector, impact, and vulnerability type in the CVE description, we verify the effectiveness of completing the key aspects of the vulnerability in predicting the severity of the security vulnerability. Zhenchang Xing, Sen Chen 0001, Xiaohong Li 0001, Yude Bai |
COMPSAC | 5 |
| 2021 | Predicting Entity Relations across Different Security Databases by Using Graph Attention NetworkabstractSecurity databases such as Common Vulnerabilities and Exposures (CVE), Common Weakness Enumeration (CWE), and Common Attack Pattern Enumeration and Classification (CAPEC) maintain diverse high-quality security concepts, which are treated as security entities. Meanwhile, security entities are documented with many potential relation types that profit for security analysis and comprehension across these three popular databases. To support reasoning security entity relationships, translation-based knowledge graph representation learning treats each triple independently for the entity prediction. However, it neglects the important semantic information about the neighbor entities around the triples. To address it, we propose a text-enhanced graph attention network model (text-enhanced GAT). This model highlights the importance of the knowledge in the 2−hop neighbors surrounding a triple, under the observation of the diversity of each entity. Thus, we can capture more structural and textual information from the knowledge graph about the security databases. Extensive experiments are designed to evaluate the effectiveness of our proposed model on the prediction of security entity relationships. Moreover, the experimental results outperform the state-of-the-art by Mean Reciprocal Rank (MRR) 0.132 for detecting the missing relationships. Yude Bai, Zhenchang Xing, Sen Chen 0001, Xiaohong Li 0001, Zhidong Deng |
COMPSAC | 2 |
| 2021 | A Character-Level Convolutional Neural Network for Predicting Exploitability of VulnerabilityabstractThe continuous discovery of software vulnerabilities have brought great challenges to the cyber security, which will lead to severe systematical or individual losses after being exploited. But the harshly increasing of software vulnerabilities overwhelms the time consuming vulnerability analysis. Security experts must pay more attention to the ones which have the highest priority to be repaired. In general, both severity and exploitability determine the severity of a software vulnerability. Compared with the severity evaluated by the Common Vulnerability Scoring System (CVSS score), the exploitability is still lack of a well-accepted standard. Furthermore, based on the perspective of attack and defense, we found that the exploitability of vulnerabilities is more attractive to hackers so that system or individual is severely affected by the exploitability rather than the severity. In this paper, we propose a deep learning based approach to predict the exploitability of the vulnerability by using the correlated textual description and characteristics. Specifically, our approach takes character-level Convolutional Neural Network (charCNN) to fetch more fine-grained character-level features from the vulnerability description instead of the word-level features considered by the previous literatures. And we highlight the importance of vulnerability characteristics such as Confidentiality Impact, Integrity Impact, Attack Vector etc. during the determination of vulnerability exploitability. Extensive experiments are set to prove the effectiveness of the given charCNN approach through the comparison on both different levels of features and different neural network models. Our approach achieves the best F1 values 93.1% (at least 2.2% more than the baselines). And we also investigate the efficiency of charCNN trained by historical vulnerability when predicting the exploitability of the newly published vulnerabilities. Finally, we further explore the robustness of the proposed model by changing the scale of training sets. For the prediction of vulnerability exploitability, we recommend to adopt 40.0% to 50.0% vulnerabilities to train a robust charCNN model. Jinghui Lyu, Yude Bai, Zhenchang Xing, Xiaohong Li 0001, Weimin Ge |
TASE | 2 |
| 2021 | Comparative analysis of feature representations and machine learning methods in Android family classification
Yude Bai, Zhenchang Xing, Duoyuan Ma, Xiaohong Li 0001, Zhiyong Feng 0002 |
Comput. Networks | 1 |
| 2020 | A Knowledge Graph-based Sensitive Feature Selection for Android Malware ClassificationabstractThe rapid increase in Android malware has brought great challenges to malware analysis. To deal with such a severe situation, it has been proposed an effective way which groups malware with common behaviors into the same malware family. Although there are many methods for malware family classification, the most critical and primary step is always the definition of sensitive behavior in an application, which will be beneficial for the later classification task. Much existing literature has manually selected sensitive features, such as permission, or even designed graph-based features via the control flow graph. They heavily depend on expert knowledge and time-consuming malware application analysis, which means it has to focus on the mal ware itself to dig out valuable security knowledge at first. However, the zooming malware overwhelms such expensive feature definition methods. To overcome such a problem, we adopt a knowledge graph-based sensitive feature selection method for Android mal ware classification. Based on the Android Developer documentation, an Android API knowledge graph is constructed at first. We can obtain not only permission but also related critical API from this graph. Note that both hyperlink relation and similarity relation are used to find out the critical API. With the knowledge graph-based sensitive features, we represent each Android malware as a boolean feature vector and send it in to a machine learning classifier for malware classification. We evaluate our proposed methods on three well-known Android malware datasets, such as Genome, Drebin, and AMD. The experimental results show that: 1) our proposed sensitive API is advantageous for malware detection; 2) API chosen by similarity relation can marginally improve performance; 3) different permission groups also make an influence for classification. Duoyuan Ma, Yude Bai, Zhenchang Xing, Lintan Sun, Xiaohong Li 0001 |
APSEC | 2 |
| 2020 | Unsuccessful story about few shot malware family classification and siamese network to the rescueabstractTo battle the ever-increasing Android malware, malware family classification, which classifies malware with common features into a malware family, has been proposed as an effective malware analysis method. Several machine-learning based approaches have been proposed for the task of malware family classification. Our study shows that malware families suffer from several data imbalance, with many families with only a small number of malware applications (referred to as few shot malware families in this work). Unfortunately, this issue has been overlooked in existing approaches. Although existing approaches achieve high classification performance at the overall level and for large malware families, our experiments show that they suffer from poor performance and generalizability for few shot malware families, and traditionally downsampling method cannot solve the problem. To address the challenge in few shot malware family classification, we propose a novel siamese-network based learning method, which allows us to train an effective MultiLayer Perceptron (MLP) network for embedding malware applications into a real-valued, continuous vector space by contrasting the malware applications from the same or different families. In the embedding space, the performance of malware family classification can be significantly improved for all scales of malware families, especially for few shot malware families, which also leads to the significant performance improvement at the overall level. Yude Bai, Zhenchang Xing, Xiaohong Li 0001, Zhiyong Feng 0002, Duoyuan Ma |
ICSE | 1 |
| 2019 | A Co-Occurrence Recommendation Model of Software Security RequirementabstractTo guarantee the quality of software, specifying security requirements (SRs) is essential for developing systems, especially for security-critical software systems. However, using security threat to determine detailed SR is quite difficult according to Common Criteria (CC), which is too confusing and technical for non-security specialists. In this paper, we propose a Co-occurrence Recommend Model (CoRM) to automatically recommend software SRs. In this model, the security threats of product are extracted from security target documents of software, in which the related security requirements are tagged. In order to establish relationships between software security threat and security requirement, semantic similarities between different security threat is calculated by Skip-thoughts Model. To evaluate our CoRM model, over 1000 security target documents of 9 types software products are exploited. The results suggest that building a CoRM model via semantic similarity is feasible and reliable. Weimin Ge, Xiaohong Li 0001, Zhiyong Feng 0002, Xiaofei Xie, Yude Bai |
TASE | 6 |