Weikai Lu

dblp:338/1147 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0003-2854-9217ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 91% Language models and text generation · 9%
Network and information security
2 papers
Security and privacy of machine learning · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › AI safety
safety alignment
1.722025
SEA: Low-Resource Safety Alignment for Multimodal Large Language Models via Synthetic Embeddings · ACL (1) 2025
SDD: Self-Degraded Defense against Malicious Fine-tuning · ACL (1) 2025
Machine learning › Trustworthy machine learning › AI safety › safety alignment
LLM safety alignment
0.912025
SDD: Self-Degraded Defense against Malicious Fine-tuning · ACL (1) 2025
Security and privacy of machine learning › large language model safety
multimodal large language model safety
0.912025
SEA: Low-Resource Safety Alignment for Multimodal Large Language Models via Synthetic Embeddings · ACL (1) 2025
Natural language and speech › Language models and text generation
large language model
0.312025
SDD: Self-Degraded Defense against Malicious Fine-tuning · ACL (1) 2025

Methods — techniques the papers use, named apart from their topics

theoretical analysis · 1.7synthetic embeddings · 1.7safety alignment · 1.7gradient-based optimization · 1.7
YearPublicationVenuePosition
2025 SDD: Self-Degraded Defense against Malicious Fine-tuning
abstract
Open-source Large Language Models (LLMs) often employ safety alignment methods to resist harmful instructions.However, recent research shows that maliciously fine-tuning these LLMs on harmful data can easily bypass these safeguards.To counter this, we theoretically uncover why malicious fine-tuning succeeds and identify potential defense strategies.Building on the theoretical analysis, we introduce the Self-Degraded Defense (SDD) framework.SDD encourages LLMs to produce high-quality but irrelevant responses to harmful prompts.When attackers attempt malicious fine-tuning, the general capability of the LLM aligned by SDD will significantly decrease, rendering it incapable of following harmful instructions.Our experimental results confirm SDD's effectiveness against such attacks.Our code is available at https://github.com/ZeroNLP/SDD.
Weikai Lu, Ziqian Zeng
ACL (1)2
2025 SEA: Low-Resource Safety Alignment for Multimodal Large Language Models via Synthetic Embeddings
abstract
Multimodal Large Language Models (MLLMs) have serious security vulnerabilities.While safety alignment using multimodal datasets consisting of text and data of additional modalities can effectively enhance MLLM's security, it is costly to construct these datasets.Existing low-resource security alignment methods, including textual alignment, have been found to struggle with the security risks posed by additional modalities.To address this, we propose Synthetic Embedding augmented safety Alignment (SEA), which optimizes embeddings of additional modality through gradient updates to expand textual datasets.This enables multimodal safety alignment training even when only textual data is available.Extensive experiments on image, video, and audio-based MLLMs demonstrate that SEA can synthesize a high-quality embedding on a single RTX3090 GPU within 24 seconds.SEA significantly improves the security of MLLMs when faced with threats from additional modalities.To assess the security risks introduced by video and audio, we also introduced a new benchmark called VA-SafetyBench.High attack success rates across multiple MLLMs validate its challenge.Our code and data will be available at https://github.com/ZeroNLP/SEA.This paper contains harmful data and modelgenerated content that can be offensive in nature.
Weikai Lu, Huiping Zhuang, Cen Chen 0002, Ziqian Zeng
ACL (1)1
2024 State-element-aware syndrome classification based on hypergraph convolutional network
Shenghua Teng, Jishun Ma, Changen Zhou, Weikai Lu
Expert Syst. Appl.5
2022 Self-supervised domain adaptation for cross-domain fault diagnosis
abstract
Unsupervised domain adaptation-based fault diagnosis methods have been extensively studied due to their powerful knowledge transferability under different working conditions. Despite their encouraging performance, most of them cannot sufficiently account for the temporal dimension of the vibration signal, resulting in incomplete feature information used in the domain alignment procedure. To alleviate the limitation, we present a self-supervised domain adaptation fault diagnosis network (SDAFDN), which considers two temporal dependencies to improve the transferability of the learned representations. Specifically, we first design a down-sampling and interaction network that considers the temporal dependency among subsequences with low temporal resolution in feature space. Then, we combine domain adversarial learning with feature mapping to achieve domain alignment. Finally, we introduced a self-supervised learning module, which considers the temporal dependency between the past and future temporal segments via classification tasks. Extensive experiments on public Paderborn University and PHM data sets demonstrate the superiority of the proposed SDAFDN and the effectiveness of considering temporal dependencies in domain alignment.
Weikai Lu, Haoyi Fan
Int. J. Intell. Syst.1