Yimo Ren

dblp:301/6126 · DBLP profile ↗
← Back
27ranked-venue papers
11as first author
27since 2021 · last 2026
0000-0002-7543-0326ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Computer networks · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SpeechShield: Latency-Efficient and Robust Timbre-Aware Voice Protection Against Speech Synthesis Deepfake Attacks
Jianshuo Liu, Shiquan Dong, Hong Li 0004, Chenghua Gao, Kang G. Shin, Haining Wang 0001, Yimo Ren, Limin Sun 0001
DSN7
2026 Breaking Cross-modal Alignment in Embodied Intelligence: A Multimodal Adversarial Attack Framework for Vision-Language-Action Models
abstract
Vision–Language–Action (VLA) models underpin robotic and other embodied agents by mapping visual observations and language instructions into executable actions. Their wide adoption through open web model repositories, however, introduces new supply-chain risks: adversaries can launch adversarial attacks to manipulate the action outputs of VLAs, potentially leading to harmful real-world outcomes for embodied agents. To exploit this vulnerability, we propose MAVLA, a novel multimodal adversarial attack framework. MAVLA serves as a modular front-end that integrates seamlessly with a target VLA model, injecting perturbations into task-relevant and structure-sensitive image regions to disrupt cross-modal alignment and induce deviations in the generated action instructions. To balance attack effectiveness with stealth, we design four loss functions that jointly maximize multimodal misalignment while preserving visual stealthiness. Extensive evaluations in simulated and real-world scenarios show that at a 40% perturbation ratio, the task success rate of VLAs drops by about 70%. Compared to conventional attack baselines, MAVLA achieves superior attack effectiveness and stealthiness with low overhead. Our work reveals a practical and previously underexplored threat to embodied systems, and offers a red-team baseline to inform future defensive strategies and promote safer VLA deployment.
Xiaorong Dong, Yaowen Zheng, Yimo Ren, Hangbei Cheng, Yongle Chen, Limin Sun 0001
WWW5
2026 StruFSM: Byte-level structural modeling for protocol finite state machine inference
Zhen Wang 0043, Yimo Ren, Zhaoteng Yan, Hong Li 0004, Hongsong Zhu
Comput. Networks4
2026 SMTFL: Secure Model Training to Untrusted Participants in Federated Learning
abstract
Federated learning is an essential distributed model training technique. However, threats such as gradient inversion attacks and poisoning attacks pose significant risks to both privacy of training data and model correctness. We propose SMTFL, a novel approach for secure model training in federated learning. To safeguard gradients privacy against gradient inversion attacks, clients are dynamically grouped, allowing one client's gradient to be divided to obfuscate the gradients of other clients within the group. This method incorporates checks and balances to reduce the collusion for inferring specific client data. To detect poisoning attacks from malicious clients, we assess the impact of aggregated gradients on the global model's performance, enabling effective identification and exclusion of malicious clients. Each client's gradients are encrypted and stored, with decryption collectively managed by all clients. The detected poisoning gradients are invalidated from the global model through an unlearning method. Compared to related work, SMTFL does not rely on trusted participants, avoids the performance degradation caused with traditional noise-injection, and avoids complex homomorphic encryption during gradient aggregation. SMTFL is evaluated on five datasets, and these results demonstrate its effectiveness in defending against gradient inversion and poisoning attacks. The model accuracy is nearly restored to its pre-attack state when SMTFL is deployed. Furthermore, SMTFL achieves over 95% accuracy in identifying malicious clients while maintaining a false positive rate for honest clients within 5%, which is 6% lower than the latest methods.
Xiaorong Dong, Yimo Ren, Jianhua Wang 0004, Hongsong Zhu, Yongle Chen
IEEE Trans. Mob. Comput.3
2025 HF-Mamba: Improving Multimodal Classification via Hierarchical Fusion Based on Mamba
Yimo Ren, Jinfa Wang, Hong Li 0004, Rongrong Xi, Haiqiang Fei, Hongsong Zhu
DASFAA (2)1
2025 Leveraging Fine-Tuned Large Language Models for Device Fingerprint Extraction in IoT Security
Haoyu Bin, Gaosheng Wang, Yimo Ren, Zhi Li 0018, Hongsong Zhu
ICIC (4)3
2025 Lares: LLM-driven Code Slice Semantic Search for Patch Presence Testing
abstract
In modern software ecosystems, 1-day vulnerabilities pose significant security risks due to extensive code reuse. Identifying vulnerable functions in target binaries alone is insufficient; it is also crucial to determine whether these functions have been patched. Existing methods, however, suffer from limited usability and accuracy. They often depend on the compilation process to extract features, requiring substantial manual effort and failing for certain software. Moreover, they cannot reliably differentiate between code changes caused by patches or compilation variations.To overcome these limitations, we propose Lares, a scalable and accurate method for patch presence testing. Lares introduces Code Slice Semantic Search, which directly extracts features from the patch source code and identifies semantically equivalent code slices in the pseudocode of the target binary. By eliminating the need for the compilation process, Lares improves usability, while leveraging large language models (LLMs) for code analysis and SMT solvers for logical reasoning to enhance accuracy. Experimental results show that Lares achieves superior precision, recall, and usability. Furthermore, it is the first work to evaluate patch presence testing across optimization levels, architectures, and compilers. The datasets and source code used in this article are available at https://github.com/Siyuan-Li201/Lares.
Siyuan Li 0014, Yaowen Zheng, Hong Li 0004, Jingdong Guo, Chaopeng Dong, Chunpeng Yan, Weijie Wang 0005, Yimo Ren, Limin Sun 0001, Hongsong Zhu
ASE8
2025 Automated Penetration on Multi-Subnet Environments with Dual-Stage DRL Models
abstract
With the advent of artificial intelligence techniques, the field of Network Attack Defense (NAD) has witnessed a surge in research efforts towards automating penetration testing (PenTest). Our work presents a dual-stage PenTest model aiming at predicting attack paths in network topology and determining payload for vulnerabilities in hosts with deep reinforcement learning models. While constructing training environments, our approach integrates real-world vulnerability environments with virtual network topologies. This allows the model to take into account the process of vulnerability validation with success rate compared to existing work based on fully virtualized targets, while retaining the efficiency of deployment and training provided by virtualization. And we introduce a method that simulate hierarchical network topology with randomized subnets to simulate complex network environments, challenging the agent to adapt and learn effective policies across diverse configurations of the target networks. Our experiments demonstrate the effectiveness of our model in various network sizes. In addition, the results indicate that our approach not only achieves high performance but also maintains stability under the different success rate of vulnerability exploitation, showcasing the robustness and adaptability. Our work contributes to the advancement of automated PenTest by providing a more generalized and efficient solution.
Haoyu Bu, Hui Wen 0001, Hongsong Zhu, Hong Li 0004, Xirui Song, Yimo Ren
SMC6
2025 MTSec: AIGC-enhanced security model training for multimodal federated learning
Xiaorong Dong, Hangbei Cheng, Yimo Ren, Yongle Chen
Knowl. Based Syst.4
2025 Automated tactics planning for cyber attack and defense based on large language model agents
Yimo Ren, Jinfa Wang, Hui Wen 0001, Hong Li 0004, Hongsong Zhu
Neural Networks1
2025 TransferFuzz-Pro: Large Language Model Driven Code Debugging Technology for Verifying Propagated Vulnerability
abstract
Code reuse in software development frequently facilitates the spread of vulnerabilities, leading to imprecise scopes of affected software in CVE reports. Traditional methods focus primarily on detecting reused vulnerability code in target software but lack the ability to confirm whether these vulnerabilities can be triggered in new software contexts. In previous work, we introduced the TransferFuzz framework to address this gap by using historical trace-based fuzzing. However, its effectiveness is constrained by the need for manual intervention and reliance on source code instrumentation. To overcome these limitations, we propose TransferFuzz-Pro, a novel framework that integrates Large Language Model (LLM)-driven code debugging technology. By leveraging LLM for automated, human-like debugging and Proof-of-Concept (PoC) generation, combined with binary-level instrumentation, TransferFuzz-Pro extends verification capabilities to a wider range of targets. Our evaluation shows that TransferFuzz-Pro is significantly faster and can automatically validate vulnerabilities that were previously unverifiable using conventional methods. Notably, it expands the number of affected software instances for 15 CVE-listed vulnerabilities from 15 to 53 and successfully generates PoCs for various Linux distributions. These results demonstrate that TransferFuzz-Pro effectively verifies vulnerabilities introduced by code reuse in target software and automatically generation PoCs.
Siyuan Li 0014, Kaiyu Xie, Yuekang Li, Hong Li 0004, Yimo Ren, Limin Sun 0001, Hongsong Zhu
IEEE Trans. Software Eng.5
2024 Hierarchical Aligned Multimodal Learning for NER on Tweet Posts
abstract
Mining structured knowledge from tweets using named entity recognition (NER) can be beneficial for many downstream applications such as recommendation and intention under standing. With tweet posts tending to be multimodal, multimodal named entity recognition (MNER) has attracted more attention. In this paper, we propose a novel approach, which can dynamically align the image and text sequence and achieve the multi-level cross-modal learning to augment textual word representation for MNER improvement. To be specific, our framework can be split into three main stages: the first stage focuses on intra-modality representation learning to derive the implicit global and local knowledge of each modality, the second evaluates the relevance between the text and its accompanying image and integrates different grained visual information based on the relevance, the third enforces semantic refinement via iterative cross-modal interactions and co-attention. We conduct experiments on two open datasets, and the results and detailed analysis demonstrate the advantage of our model.
Hong Li 0004, Yimo Ren, Jie Liu 0079, Shuaizong Si, Hongsong Zhu, Limin Sun 0001
AAAI3
2024 A Relation-Aware Heterogeneous Graph Transformer on Dynamic Fusion for Multimodal Classification Tasks
abstract
Multimodal fusion aims to improve the performance of models for applications by extracting and fusing information in different modalities, including texts, images or others. Recent researches have shown that multimodal fusion is beneficial in many multimedia tasks. In this paper, we study typical multimedia classification tasks in social media posts, including sarcasm detection and sentiment analysis. This paper proposes DMF-RHGT-HPA, including dynamic Fusion multimodal fusion(DMF), a relation-aware heterogeneous graph transformer(RHGT) and hierarchical pooling alignment(HPA). To realize better multimodal fusion, the paper designs it on a heterogeneous graph with dynamic links, without any padding of texts or images. To thoroughly learn the multimodal graph and obtain the representation of nodes, the paper proposes a relation-aware heterogeneous graph transformer to fuse the node-level and edge-level features simultaneously. To get a refined representation of the multimodal graph, the paper designs a hierarchical pooling alignment to gather all nodes’ representations well. Experiments conducted on two primary and public datasets from Twitter and Yelp respectively show the ability of DMF-RHGT-HPA to gain the best performance of sarcasm detection and sentiment analysis, outperforming existing state-of-the-art baselines.
Yimo Ren, Jinfa Wang, Jie Liu 0079, Hong Li 0004, Hongsong Zhu, Limin Sun 0001
ICASSP1
2024 Lexicon Graph Adapter Based BERT Model for Chinese Named Entity Recognition
Jie Liu 0079, Yimo Ren, Jinfa Wang, Hongsong Zhu
KSEM (5)3
2024 Multi-granularity cross-modal representation learning for named entity recognition on social media
Gaosheng Wang, Hong Li 0004, Jie Liu 0079, Yimo Ren, Hongsong Zhu, Limin Sun 0001
Inf. Process. Manag.5
2023 Improving the Modality Representation with multi-view Contrastive Learning for Multimodal Sentiment Analysis
abstract
Modality representation learning is an important problem for multimodal sentiment analysis (MSA), since the highly distinguishable representations can contribute to improving the analysis effect. Previous works of MSA have usually focused on internal fusion strategies for different modalities within one sample, and the external usage of cross reference relations among different samples was given less attention. Recently, the rise of contrastive learning provides powerful clues for us to learn modal representation with stronger discriminative ability. In this study, we explore the approach of representations improvement and devise a three-stages framework with multi-view contrastive learning to refine representations for the specific objectives. Firstly, for each modality, we employ the supervised contrastive learning to pull samples within the same class together while the other samples are pushed apart. Then, a self-supervised contrastive learning is designed for the distilled cross-modal representations after a novel Transformer-based interaction module. At last, we leverage again the supervised contrastive learning to enhance the fused multimodal representation. We conduct extensive experiments on three open datasets, and results show the advance of our model.
Hong Li 0004, Jie Liu 0079, Yimo Ren, Hongsong Zhu, Limin Sun 0001
ICASSP5
2023 CEntRE: A paragraph-level Chinese dataset for Relation Extraction among Enterprises
abstract
Enterprise relation extraction aims to detect pairs of enterprise entities and identify the business relations between them from unstructured or semi-structured text data, and it is crucial for several real-world applications such as risk analysis, rating research and supply chain security. However, previous work mainly focuses on getting attribute information about enterprises like personnel and corporate business, and pays little attention to enterprise relation extraction. To encourage further progress in the research, we introduce the CEntRE, a new dataset constructed from publicly available business news data with careful human annotation and intelligent data processing. Moreover, we propose a joint entity and relation extraction network, which is capable of discovering enterprise entities and extracting business relations between them accurately. The network firstly encodes input sequences with strong semantic augmentation to learn contextual representation for each token, then a conditional random field (CRF) module is used for entity extraction. Subsequently, entity pairs are built and a new encoder based on the entity pairs is applied to get global information for relation extraction. Finally, a biaffine classifier is deployed to classify the relations. Extensive experiments on CEntRE demonstrate the effectiveness of our proposed method compared with other six excellent models, and thus our model can be considered as one strong baseline. The data and code are available at: https://github.com/LiuPeiP-CStMining_Entity_Relations_Among_Enterprises
Hong Li 0004, Yimo Ren, Jie Liu 0079, Fei Lyu 0001, Hongsong Zhu, Limin Sun 0001
IJCNN4
2023 User Recognition of Devices on the Internet based on Heterogeneous Graph Transformer with Partial Labels
abstract
Recognizing the users of devices can easily enable numerous security applications. Due to the lot's kinds of device data and a large number of missing values, it takes work to recognize the users of devices well. The community detection methods based on Graph Neural Networks (GNN) can integrate multi-source data well and cluster devices into communities with the same users. While existing GNN methods face several issues. The methods on homogeneous graphs could not utilize the multi-source data of devices, and most methods on heterogeneous graphs need specific knowledge to design meta paths. Also, the Internet-scale data of devices make it hard to learn the representation thoroughly. Further, most methods need to consider the known partial labels in the early stage of the training process. To improve the performance of user recognition, this paper proposes HGT-PL, namely a Heterogeneous Graph Transformer with Partial Labels, to calculate the representation of devices on the Internet. Then cluster methods are used to realize user recognition. Using graph transformers, HGT-PL deeply learns node features and graph structure on the heterogeneous graph of devices. By Label Encoder, HGT-PL fully utilizes the users of partial devices from preliminary rules with high confidence. Moreover, cluster methods carefully divide and modify the communities with different users. The paper conducts experiments on the web-scale data collected from the Internet. The results show that HGT-PL can recognize users of devices more accurately and effectively, with 0.5121 NMI and 0.3554 ARI, compared with existing GNN methods.
Yimo Ren, Jinfa Wang, Hong Li 0004, Hongsong Zhu, Limin Sun 0001
IJCNN1
2023 DeviceGPT: A Generative Pre-Training Transformer on the Heterogenous Graph for Internet of Things
abstract
Recently, Graph neural networks (GNNs) have been adopted to model a wide range of structured data from academic and industry fields. With the rapid development of Internet technology, there are more and more meaningful applications for Internet devices, including device identification, geolocation and others, whose performance needs improvement. To replicate the several claimed successes of GNNs, this paper proposes DeviceGPT based on a generative pre-training transformer on a heterogeneous graph via self-supervised learning to learn interactions-rich information of devices from its large-scale databases well. The experiments on the dataset constructed from the real world show DeviceGPT could achieve competitive results in multiple Internet applications.
Yimo Ren, Jinfa Wang, Hong Li 0004, Hongsong Zhu, Limin Sun 0001
SIGIR1
2023 CL-GAN: A GAN-based continual learning model for generating and detecting AGDs
Yimo Ren, Hong Li 0004, Jie Liu 0079, Hongsong Zhu, Limin Sun 0001
Comput. Secur.1
2023 Owner name entity recognition in websites based on multiscale features and multimodal co-attention
Yimo Ren, Hong Li 0004, Jie Liu 0079, Hongsong Zhu, Limin Sun 0001
Expert Syst. Appl.1
2023 Multiview Embedding with Partial Labels to Recognize Users of Devices Based on Unified Transformer
abstract
Recognizing the users of devices (or clusters of devices) who use IP addresses as unique identities on the Internet can easily enable numerous security applications. Fast and accurate user recognition is critical for supervisors to find influenced organizations connected to their networks in light of new security threats. Many users’ information scatters in the multisource data of IP addresses. Up until now, user recognition of devices has had two main problems. On the one hand, existing methods could not fully use multisource data of the IP addresses and wastes the valuable information of labels. On the other hand, only a tiny portion of devices can be tagged with highly confident known users manually, making it an urgent need to infer unknown users of devices. So, the problem of user recognition on devices is to guess the unknown user with multisource data and existing devices with known users. Therefore, this paper proposes a multiview fusion method to deal with multisource data from devices with a small number of manually labelled samples. The paper uses GraphSAGE to obtain an exemplary representation of IP addresses and designs a label encoder to fully use a small number of devices with known users. Then, the paper builds a specific unified transformer to achieve high performance to determine whether two devices have the same user. At the same time, the paper conducts real‐world experiments and finds that the proposed method can achieve 0.9158 accuracy and 0.6131 F1 to find devices with the same users on the constructed dataset in the real world.
Yimo Ren, Hong Li 0004, Jie Liu 0079, Hongsong Zhu, Limin Sun 0001
Int. J. Intell. Syst.1
2023 Owner name entity recognition in websites based on heterogeneous and dynamic graph transformer
Yimo Ren, Hong Li 0004, Jie Liu 0079, Zhi Li 0018, Hongsong Zhu, Limin Sun 0001
Knowl. Inf. Syst.1
2023 Owner named entity recognition in website based on multidimensional text guidance and space alignment co-attention
Xin He 0021, Yimo Ren, Jinfa Wang, Junyang Yu
Multim. Syst.3
2022 Multi-features based Semantic Augmentation Networks for Named Entity Recognition in Threat Intelligence
abstract
Extracting cybersecurity entities such as attackers and vulnerabilities from unstructured network texts is an important part of security analysis. However, the sparsity of intelligence data resulted from the higher frequency variations and the randomness of cybersecurity entity names makes it difficult for current methods to perform well in extracting security-related concepts and entities. To this end, we propose a semantic augmentation method which incorporates different linguistic features to enrich the representation of input tokens to detect and classify the cybersecurity names over unstructured text. In particular, we encode and aggregate the constituent feature, morphological feature and part of speech feature for each input token to improve the robustness of the method. More than that, a token gets augmented semantic information from its most similar K words in cybersecurity domain corpus where an attentive module is leveraged to weigh differences of the words, and from contextual clues based on a large-scale general field corpus. We have conducted experiments on the cybersecurity datasets DNRTI and MalwareTextDB, and the results demonstrate the effectiveness of the proposed method.
Hong Li 0004, Zuoguang Wang, Jie Liu 0079, Yimo Ren, Hongsong Zhu
ICPR5
2022 Joint Classification of IoT Devices and Relations in the Internet with Network Traffic
abstract
With the rapid growth and popularization of Internet of Things (IoT), more and more devices are deployed in homes, enterprises, cities, etc. The existed methods to classify types and relations of devices are usually two separate tasks. So, it is difficult to quickly provide attributes of devices in the smart network for operators at the same time. At this situation, the paper presents a framework JCIDR for Joint Classification of IoT Device and Relations In the Internet with Network Traffic. By fusing the numerical features and binary image features of traffic, the devices and relations of devices can be recognized simultaneously. The experiment is carried out in a real IoT environment and the accuracy of JCIDR is over 86% with about half time reduction. Therefore, JCIDR could provide operators with a fast, easy, low-cost network device monitoring method without professional equipment or protocols.
Yimo Ren, Hong Li 0004, Shuqin Zhang, Hongsong Zhu, Limin Sun 0001
WCNC1
2021 FIUD: A Framework to Identify Users of Devices
Yimo Ren, Hong Li 0004, Hongsong Zhu, Limin Sun 0001
WASA (2)1