Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Runqi Lin

dblp:359/1108 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Trustworthy machine learning · 64% Deep learning architectures and training · 15% Learning theory · 13%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
2.232024
Layer-Aware Analysis of Catastrophic Overfitting: Revealing the Pseudo-Robust Shortcut Dependency · ICML 2024
On the Over-Memorization During Natural, Robust and Catastrophic Overfitting · ICLR 2024
Eliminating Catastrophic Overfitting Via Abnormal Adversarial Examples Regularization · NeurIPS 2023
Machine learning › Trustworthy machine learning
robustness
2.232024
Layer-Aware Analysis of Catastrophic Overfitting: Revealing the Pseudo-Robust Shortcut Dependency · ICML 2024
On the Over-Memorization During Natural, Robust and Catastrophic Overfitting · ICLR 2024
Eliminating Catastrophic Overfitting Via Abnormal Adversarial Examples Regularization · NeurIPS 2023
Machine learning › Trustworthy machine learning › robustness › adversarial robustness › adversarial training › fast adversarial training
catastrophic overfitting
1.422024
Layer-Aware Analysis of Catastrophic Overfitting: Revealing the Pseudo-Robust Shortcut Dependency · ICML 2024
Eliminating Catastrophic Overfitting Via Abnormal Adversarial Examples Regularization · NeurIPS 2023
Machine learning › Deep learning architectures and training › regularization
early stopping
0.912025
Instance-dependent Early Stopping · ICLR 2025
Machine learning › Deep learning architectures and training › regularization
training regularization
0.912025
Instance-dependent Early Stopping · ICLR 2025
Security and privacy of machine learning › adversarial attack
jailbreak attack
0.912025
Understanding and Enhancing the Transferability of Jailbreaking Attacks · ICLR 2025
Machine learning › Trustworthy machine learning › robustness › adversarial attack
adversarial weight perturbation
0.812024
Layer-Aware Analysis of Catastrophic Overfitting: Revealing the Pseudo-Robust Shortcut Dependency · ICML 2024
Machine learning › Learning theory
generalization
0.812024
On the Over-Memorization During Natural, Robust and Catastrophic Overfitting · ICLR 2024
Machine learning › Learning theory
overfitting
0.812024
On the Over-Memorization During Natural, Robust and Catastrophic Overfitting · ICLR 2024
Machine learning › Trustworthy machine learning › robustness › adversarial robustness › adversarial training
single-step adversarial training
0.712023
Eliminating Catastrophic Overfitting Via Abnormal Adversarial Examples Regularization · NeurIPS 2023
Machine learning › Trustworthy machine learning › robustness
adversarial examples
0.212023
Eliminating Catastrophic Overfitting Via Abnormal Adversarial Examples Regularization · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

second-order loss differences · 0.9red-teaming evaluation · 0.9backpropagation exclusion · 0.9adversarial sequence optimization · 0.9weight perturbation · 0.8distraction over-memorization · 0.8adversarial training · 0.8inner maximization · 0.7abnormal adversarial examples regularization · 0.7
YearPublicationVenuePosition
2025 Understanding and Enhancing the Transferability of Jailbreaking Attacks
abstract
Jailbreaking attacks can effectively manipulate open-source large language models (LLMs) to produce harmful responses. However, these attacks exhibit limited transferability, failing to disrupt proprietary LLMs consistently. To reliably identify vulnerabilities in proprietary LLMs, this work investigates the transferability of jailbreaking attacks by analysing their impact on the model's intent perception. By incorporating adversarial sequences, these attacks can redirect the source LLM's focus away from malicious-intent tokens in the original input, thereby obstructing the model's intent recognition and eliciting harmful responses. Nevertheless, these adversarial sequences fail to mislead the target LLM's intent perception, allowing the target LLM to refocus on malicious-intent tokens and abstain from responding. Our analysis further reveals the inherent $\textit{distributional dependency}$ within the generated adversarial sequences, whose effectiveness stems from overfitting the source LLM's parameters, resulting in limited transferability to target LLMs. To this end, we propose the Perceived-importance Flatten (PiF) method, which uniformly disperses the model's focus across neutral-intent tokens in the original input, thus obscuring malicious-intent tokens without relying on overfitted adversarial sequences. Extensive experiments demonstrate that PiF provides an effective and efficient red-teaming evaluation for proprietary LLMs.
Runqi Lin, Bo Han 0003, Fengwang Li, Tongliang Liu
ICLR1
2025 Instance-dependent Early Stopping
abstract
In machine learning practice, early stopping has been widely used to regularize models and can save computational costs by halting the training process when the model's performance on a validation set stops improving. However, conventional early stopping applies the same stopping criterion to all instances without considering their individual learning statuses, which leads to redundant computations on instances that are already well-learned. To further improve the efficiency, we propose an Instance-dependent Early Stopping (IES) method that adapts the early stopping mechanism from the entire training set to the instance level, based on the core principle that once the model has mastered an instance, the training on it should stop. IES considers an instance as mastered if the second-order differences of its loss value remain within a small range around zero. This offers a more consistent measure of an instance's learning status compared with directly using the loss value, and thus allows for a unified threshold to determine when an instance can be excluded from further backpropagation. We show that excluding mastered instances from backpropagation can increase the gradient norms, thereby accelerating the decrease of the training loss and speeding up the training process. Extensive experiments on benchmarks demonstrate that IES method can reduce backpropagation instances by 10%-50% while maintaining or even slightly improving the test accuracy and transfer learning performance of a model.
Suqin Yuan, Runqi Lin, Lei Feng 0006, Bo Han 0003, Tongliang Liu
ICLR2
2024 On the Over-Memorization During Natural, Robust and Catastrophic Overfitting
abstract
Overfitting negatively impacts the generalization ability of deep neural networks (DNNs) in both natural and adversarial training. Existing methods struggle to consistently address different types of overfitting, typically designing strategies that focus separately on either natural or adversarial patterns. In this work, we adopt a unified perspective by solely focusing on natural patterns to explore different types of overfitting. Specifically, we examine the memorization effect in DNNs and reveal a shared behaviour termed over-memorization, which impairs their generalization capacity. This behaviour manifests as DNNs suddenly becoming high-confidence in predicting certain training patterns and retaining a persistent memory for them. Furthermore, when DNNs over-memorize an adversarial pattern, they tend to simultaneously exhibit high-confidence prediction for the corresponding natural pattern. These findings motivate us to holistically mitigate different types of overfitting by hindering the DNNs from over-memorization training patterns. To this end, we propose a general framework, $\textit{Distraction Over-Memorization}$ (DOM), which explicitly prevents over-memorization by either removing or augmenting the high-confidence natural patterns. Extensive experiments demonstrate the effectiveness of our proposed method in mitigating overfitting across various training paradigms.
Runqi Lin, Chaojian Yu, Bo Han 0003, Tongliang Liu
ICLR1
2024 Layer-Aware Analysis of Catastrophic Overfitting: Revealing the Pseudo-Robust Shortcut Dependency
abstract
Catastrophic overfitting (CO) presents a significant challenge in single-step adversarial training (AT), manifesting as highly distorted deep neural networks (DNNs) that are vulnerable to multi-step adversarial attacks. However, the underlying factors that lead to the distortion of decision boundaries remain unclear. In this work, we delve into the specific changes within different DNN layers and discover that during CO, the former layers are more susceptible, experiencing earlier and greater distortion, while the latter layers show relative insensitivity. Our analysis further reveals that this increased sensitivity in former layers stems from the formation of $\textit{pseudo-robust shortcuts}$, which alone can impeccably defend against single-step adversarial attacks but bypass genuine-robust learning, resulting in distorted decision boundaries. Eliminating these shortcuts can partially restore robustness in DNNs from the CO state, thereby verifying that dependence on them triggers the occurrence of CO. This understanding motivates us to implement adaptive weight perturbations across different layers to hinder the generation of $\textit{pseudo-robust shortcuts}$, consequently mitigating CO. Extensive experiments demonstrate that our proposed method, $\textbf{L}$ayer-$\textbf{A}$ware Adversarial Weight $\textbf{P}$erturbation (LAP), can effectively prevent CO and further enhance robustness.
Runqi Lin, Chaojian Yu, Bo Han 0003, Hang Su 0006, Tongliang Liu
ICML1
2023 Eliminating Catastrophic Overfitting Via Abnormal Adversarial Examples Regularization
abstract
Single-step adversarial training (SSAT) has demonstrated the potential to achieve both efficiency and robustness. However, SSAT suffers from catastrophic overfitting (CO), a phenomenon that leads to a severely distorted classifier, making it vulnerable to multi-step adversarial attacks. In this work, we observe that some adversarial examples generated on the SSAT-trained network exhibit anomalous behaviour, that is, although these training samples are generated by the inner maximization process, their associated loss decreases instead, which we named abnormal adversarial examples (AAEs). Upon further analysis, we discover a close relationship between AAEs and classifier distortion, as both the number and outputs of AAEs undergo a significant variation with the onset of CO. Given this observation, we re-examine the SSAT process and uncover that before the occurrence of CO, the classifier already displayed a slight distortion, indicated by the presence of few AAEs. Furthermore, the classifier directly optimizing these AAEs will accelerate its distortion, and correspondingly, the variation of AAEs will sharply increase as a result. In such a vicious circle, the classifier rapidly becomes highly distorted and manifests as CO within a few iterations. These observations motivate us to eliminate CO by hindering the generation of AAEs. Specifically, we design a novel method, termed Abnormal Adversarial Examples Regularization (AAER), which explicitly regularizes the variation of AAEs to hinder the classifier from becoming distorted. Extensive experiments demonstrate that our method can effectively eliminate CO and further boost adversarial robustness with negligible additional computational overhead. Our implementation can be found at https://github.com/tmllab/2023_NeurIPS_AAER.
Runqi Lin, Chaojian Yu, Tongliang Liu
NeurIPS1