David Dobre

dblp:322/2336 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Optimization for machine learning · 34% Trustworthy machine learning · 29% Generative modeling · 22%
Network and information security
2 papers
Security and privacy of machine learning · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › large language model safety
safety fine-tuning
0.912025
Learning Diverse Attacks on Large Language Models for Robust Red-Teaming and Safety Tuning · ICLR 2025
Security and privacy of machine learning
red teaming
0.912025
Learning Diverse Attacks on Large Language Models for Robust Red-Teaming and Safety Tuning · ICLR 2025
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.812024
On the Scalability of Certified Adversarial Robustness with Generated Data · NeurIPS 2024
Machine learning › Trustworthy machine learning › robustness
certified robustness
0.812024
On the Scalability of Certified Adversarial Robustness with Generated Data · NeurIPS 2024
Machine learning › Generative modeling
diffusion model
0.812024
On the Scalability of Certified Adversarial Robustness with Generated Data · NeurIPS 2024
Machine learning › Generative modeling
synthetic training data
0.812024
On the Scalability of Certified Adversarial Robustness with Generated Data · NeurIPS 2024
Security and privacy of machine learning
adversarial attack
0.812024
Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space · NeurIPS 2024
Mathematical optimization
frank-wolfe algorithm
0.812024
Sarah Frank-Wolfe: Methods for Constrained Optimization with Best Rates and Practical Features · ICML 2024
Mathematical optimization
stochastic optimization
0.812024
Sarah Frank-Wolfe: Methods for Constrained Optimization with Best Rates and Practical Features · ICML 2024
Machine learning › Optimization for machine learning › stochastic optimization
heavy-tailed noise
0.612022
Clipped Stochastic Methods for Variational Inequalities with Heavy-Tailed Noise · NeurIPS 2022
Machine learning › Optimization for machine learning
minimax optimization
0.612022
Clipped Stochastic Methods for Variational Inequalities with Heavy-Tailed Noise · NeurIPS 2022
Machine learning › Optimization for machine learning
stochastic optimization
0.612022
Clipped Stochastic Methods for Variational Inequalities with Heavy-Tailed Noise · NeurIPS 2022
Machine learning › Optimization for machine learning
variational inequality
0.612022
Clipped Stochastic Methods for Variational Inequalities with Heavy-Tailed Noise · NeurIPS 2022
Computer vision › Image recognition and object detection
image classification
0.212024
On the Scalability of Certified Adversarial Robustness with Generated Data · NeurIPS 2024
Machine learning › Trustworthy machine learning › adversarial machine learning
jailbreak attack
0.212024
Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space · NeurIPS 2024
Machine learning › Trustworthy machine learning › AI safety › safety alignment
LLM safety alignment
0.212024
Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space · NeurIPS 2024
Machine learning › Trustworthy machine learning
machine unlearning
0.212024
Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space · NeurIPS 2024
Machine learning › Generative modeling › generative adversarial network
GAN training
0.212022
Clipped Stochastic Methods for Variational Inequalities with Heavy-Tailed Noise · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.7stochastic finite-sum minimization · 1.5projection-free optimization · 1.5frank-wolfe · 1.5discrete jailbreak transfer · 1.5continuous embedding perturbation · 1.5GFlowNets · 0.9GFlowNet · 0.9diffusion model · 0.8adversarial training · 0.8gradient clipping · 0.6
YearPublicationVenuePosition
2025 Learning Diverse Attacks on Large Language Models for Robust Red-Teaming and Safety Tuning
abstract
Red-teaming, or identifying prompts that elicit harmful responses, is a critical step in ensuring the safe and responsible deployment of large language models (LLMs). Developing effective protection against many modes of attack prompts requires discovering diverse attacks. Automated red-teaming typically uses reinforcement learning to fine-tune an attacker language model to generate prompts that elicit undesirable responses from a target LLM, as measured, for example, by an auxiliary toxicity classifier. We show that even with explicit regularization to favor novelty and diversity, existing approaches suffer from mode collapse or fail to generate effective attacks. As a flexible and probabilistically principled alternative, we propose to use GFlowNet fine-tuning, followed by a secondary smoothing phase, to train the attacker model to generate *diverse* and *effective* attack prompts. We find that the attacks generated by our method are effective against a wide range of target LLMs, both with and without safety tuning, and transfer well between target LLMs. Finally, we demonstrate that models safety-tuned using a dataset of red-teaming prompts generated by our method are robust to attacks from other RL-based red-teaming approaches.
Seanie Lee, Minsu Kim 0004, Lynn Cherif, David Dobre, Juho Lee 0001, Sung Ju Hwang, Kenji Kawaguchi, Gauthier Gidel, Yoshua Bengio, Nikolay Malkin, Moksh Jain
ICLR4
2024 Sarah Frank-Wolfe: Methods for Constrained Optimization with Best Rates and Practical Features
abstract
The Frank-Wolfe (FW) method is a popular approach for solving optimization problems with structured constraints that arise in machine learning applications. In recent years, stochastic versions of FW have gained popularity, motivated by large datasets for which the computation of the full gradient is prohibitively expensive. In this paper, we present two new variants of the FW algorithms for stochastic finite-sum minimization. Our algorithms have the best convergence guarantees of existing stochastic FW approaches for both convex and non-convex objective functions. Our methods do not have the issue of permanently collecting large batches, which is common to many stochastic projection-free approaches. Moreover, our second approach does not require either large batches or full deterministic gradients, which is a typical weakness of many techniques for finite-sum problems. The faster theoretical rates of our approaches are confirmed experimentally.
Aleksandr Beznosikov, David Dobre, Gauthier Gidel
ICML2
2024 On the Scalability of Certified Adversarial Robustness with Generated Data
abstract
Certified defenses against adversarial attacks offer formal guarantees on the robustness of a model, making them more reliable than empirical methods such as adversarial training, whose effectiveness is often later reduced by unseen attacks. Still, the limited certified robustness that is currently achievable has been a bottleneck for their practical adoption. Gowal et al. and Wang et al. have shown that generating additional training data using state-of-the-art diffusion models can considerably improve the robustness of adversarial training. In this work, we demonstrate that a similar approach can substantially improve deterministic certified defenses but also reveal notable differences in the scaling behavior between certified and empirical methods. In addition, we provide a list of recommendations to scale the robustness of certified training approaches. Our approach achieves state-of-the-art deterministic robustness certificates on CIFAR-10 for the $\ell_2$ ($\epsilon = 36/255$) and $\ell_{\infty}$ ($\epsilon = 8/255$) threat models, outperforming the previous results by $+3.95$ and $+1.39$ percentage points, respectively. Furthermore, we report similar improvements for CIFAR-100.
Thomas Altstidl, David Dobre, Arthur Kosmala, Björn M. Eskofier, Gauthier Gidel, Leo Schwinn
NeurIPS2
2024 Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space
abstract
Current research in adversarial robustness of LLMs focuses on \textit{discrete} input manipulations in the natural language space, which can be directly transferred to \textit{closed-source} models. However, this approach neglects the steady progression of \textit{open-source} models. As open-source models advance in capability, ensuring their safety becomes increasingly imperative. Yet, attacks tailored to open-source LLMs that exploit full model access remain largely unexplored. We address this research gap and propose the \textit{embedding space attack}, which directly attacks the \textit{continuous} embedding representation of input tokens. We find that embedding space attacks circumvent model alignments and trigger harmful behaviors more efficiently than discrete attacks or model fine-tuning. Additionally, we demonstrate that models compromised by embedding attacks can be used to create discrete jailbreaks in natural language. Lastly, we present a novel threat model in the context of unlearning and show that embedding space attacks can extract supposedly deleted information from unlearned LLMs across multiple datasets and models. Our findings highlight embedding space attacks as an important threat model in open-source LLMs.
Leo Schwinn, David Dobre, Sophie Xhonneux, Gauthier Gidel, Stephan Günnemann
NeurIPS2
2022 Clipped Stochastic Methods for Variational Inequalities with Heavy-Tailed Noise
abstract
Stochastic first-order methods such as Stochastic Extragradient (SEG) or Stochastic Gradient Descent-Ascent (SGDA) for solving smooth minimax problems and, more generally, variational inequality problems (VIP) have been gaining a lot of attention in recent years due to the growing popularity of adversarial formulations in machine learning. While high-probability convergence bounds are known to more accurately reflect the actual behavior of stochastic methods, most convergence results are provided in expectation. Moreover, the only known high-probability complexity results have been derived under restrictive sub-Gaussian (light-tailed) noise and bounded domain assumptions [Juditsky et al., 2011]. In this work, we prove the first high-probability complexity results with logarithmic dependence on the confidence level for stochastic methods for solving monotone and structured non-monotone VIPs with non-sub-Gaussian (heavy-tailed) noise and unbounded domains. In the monotone case, our results match the best known ones in the light-tails case [Juditsky et al., 2011], and are novel for structured non-monotone problems such as negative comonotone, quasi-strongly monotone, and/or star-cocoercive ones. We achieve these results by studying SEG and SGDA with clipping. In addition, we numerically validate that the gradient noise of many practical GAN formulations is heavy-tailed and show that clipping improves the performance of SEG/SGDA.
Eduard Gorbunov, Marina Danilova, David Dobre, Pavel E. Dvurechensky, Alexander V. Gasnikov, Gauthier Gidel
NeurIPS3