Klas Leino

dblp:215/3626 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
8since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Trustworthy machine learning · 76% Deep learning architectures and training · 9% Kernel, tree and ensemble methods · 6%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 21 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
robustness
3.462023
Unlocking Deterministic Robustness Certification on ImageNet · NeurIPS 2023
On the Perils of Cascading Robust Classifiers · ICLR 2023
Selective Ensembles for Consistent Predictions · ICLR 2022
Machine learning › Trustworthy machine learning › robustness
certified robustness
2.952024
A Recipe for Improved Certifiable Robustness · ICLR 2024
Unlocking Deterministic Robustness Certification on ImageNet · NeurIPS 2023
Relaxing Local Robustness · NeurIPS 2021
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
2.552024
A Recipe for Improved Certifiable Robustness · ICLR 2024
Relaxing Local Robustness · NeurIPS 2021
Machine Learning Explainability and Robustness: Connected at the Hip · KDD 2021
Machine learning › Trustworthy machine learning
interpretability
0.922021
Machine Learning Explainability and Robustness: Connected at the Hip · KDD 2021
Influence Paths for Characterizing Subject-Verb Number Agreement in LSTM Language Models · ACL 2020
Machine learning › Trustworthy machine learning › verification
lipschitz certification
0.812024
A Recipe for Improved Certifiable Robustness · ICLR 2024
Computer vision › Image recognition and object detection › object detection
cascade classifier
0.712023
On the Perils of Cascading Robust Classifiers · ICLR 2023
Machine learning › Trustworthy machine learning
certifiable training
0.712023
Unlocking Deterministic Robustness Certification on ImageNet · NeurIPS 2023
Machine learning › Deep learning architectures and training › convolutional neural network
residual network
0.712023
Unlocking Deterministic Robustness Certification on ImageNet · NeurIPS 2023
Machine learning › Trustworthy machine learning › robustness
robust learning
0.712023
Unlocking Deterministic Robustness Certification on ImageNet · NeurIPS 2023
Machine learning › Kernel, tree and ensemble methods
ensemble learning
0.612022
Selective Ensembles for Consistent Predictions · ICLR 2022
Machine learning › Kernel, tree and ensemble methods › ensemble learning
ensemble pruning
0.612022
Selective Ensembles for Consistent Predictions · ICLR 2022
Machine learning › Trustworthy machine learning › robustness
prediction consistency
0.612022
Selective Ensembles for Consistent Predictions · ICLR 2022
Machine learning › Trustworthy machine learning › interpretability
post-hoc explanation
0.512021
Machine Learning Explainability and Robustness: Connected at the Hip · KDD 2021
Natural language and speech › Language models and text generation › natural language understanding › linguistic knowledge in language models
syntactic agreement
0.412020
Influence Paths for Characterizing Subject-Verb Number Agreement in LSTM Language Models · ACL 2020
Security and privacy of machine learning
membership inference
0.412020
Stolen Memories: Leveraging Model Memorization for Calibrated White-Box Membership Inference · USENIX Security Symposium 2020
Machine learning › Trustworthy machine learning › fairness › algorithmic bias
bias amplification
0.412019
Feature-Wise Bias Amplification · ICLR (Poster) 2019
Machine learning › Trustworthy machine learning
fairness
0.412019
Feature-Wise Bias Amplification · ICLR (Poster) 2019
Machine learning › Generative modeling
diffusion model
0.212024
A Recipe for Improved Certifiable Robustness · ICLR 2024
Machine learning › Deep learning architectures and training › data augmentation
generative data augmentation
0.212024
A Recipe for Improved Certifiable Robustness · ICLR 2024
Machine learning › Trustworthy machine learning › interpretability › explanation evaluation
explanation faithfulness
0.112021
Machine Learning Explainability and Robustness: Connected at the Hip · KDD 2021
Natural language and speech › Language models and text generation › large language model › knowledge in language models
memorization
0.112020
Stolen Memories: Leveraging Model Memorization for Calibrated White-Box Membership Inference · USENIX Security Symposium 2020

Methods — techniques the papers use, named apart from their topics

lipschitz constraint · 0.8filtered generative data augmentation · 0.8cholesky orthogonalization · 0.8margin maximization · 0.7lipschitz bound · 0.7cascading · 0.7ensemble learning · 0.6local robustness certification · 0.5lipschitz constant estimation · 0.5geometric projection · 0.5shadow model · 0.4calibration · 0.4
YearPublicationVenuePosition
2024 A Recipe for Improved Certifiable Robustness
abstract
Recent studies have highlighted the potential of Lipschitz-based methods for training certifiably robust neural networks against adversarial attacks. A key challenge, supported both theoretically and empirically, is that robustness demands greater network capacity and more data than standard training. However, effectively adding capacity under stringent Lipschitz constraints has proven more difficult than it may seem, evident by the fact that state-of-the-art approach tend more towards \emph{underfitting} than overfitting. Moreover, we posit that a lack of careful exploration of the design space for Lipshitz-based approaches has left potential performance gains on the table. In this work, we provide a more comprehensive evaluation to better uncover the potential of Lipschitz-based certification methods. Using a combination of novel techniques, design optimizations, and synthesis of prior work, we are able to significantly improve the state-of-the-art VRA for deterministic certification on a variety of benchmark datasets, and over a range of perturbation sizes. Of particular note, we discover that the addition of large ``Cholesky-orthogonalized residual dense'' layers to the end of existing state-of-the-art Lipschitz-controlled ResNet architectures is especially effective for increasing network capacity and performance. Combined with filtered generative data augmentation, our final results further the state of the art deterministic VRA by up to 8.5 percentage points.
Klas Leino, Zifan Wang 0001, Matt Fredrikson
ICLR2
2023 On the Perils of Cascading Robust Classifiers
Ravi Mangal, Zifan Wang 0001, Klas Leino, Corina Pasareanu, Matt Fredrikson
ICLR4
2023 Unlocking Deterministic Robustness Certification on ImageNet
abstract
Despite the promise of Lipschitz-based methods for provably-robust deep learning with deterministic guarantees, current state-of-the-art results are limited to feed-forward Convolutional Networks (ConvNets) on low-dimensional data, such as CIFAR-10. This paper investigates strategies for expanding certifiably robust training to larger, deeper models. A key challenge in certifying deep networks is efficient calculation of the Lipschitz bound for residual blocks found in ResNet and ViT architectures. We show that fast ways of bounding the Lipschitz constant for conventional ResNets are loose, and show how to address this by designing a new residual block, leading to the *Linear ResNet* (LiResNet) architecture. We then introduce *Efficient Margin MAximization* (EMMA), a loss function that stabilizes robust training by penalizing worst-case adversarial examples from multiple classes simultaneously. Together, these contributions yield new *state-of-the-art* robust accuracy on CIFAR-10/100 and Tiny-ImageNet under $\ell_2$ perturbations. Moreover, for the first time, we are able to scale up fast deterministic robustness guarantees to ImageNet, demonstrating that this approach to robust learning can be applied to real-world applications.
Andy Zou, Zifan Wang 0001, Klas Leino, Matt Fredrikson
NeurIPS4
2022 Selective Ensembles for Consistent Predictions
Emily Black, Klas Leino, Matt Fredrikson
ICLR2
2021 Fast Geometric Projections for Local Robustness Certification
Aymeric Fromherz, Klas Leino, Matt Fredrikson, Bryan Parno, Corina Pasareanu
ICLR2
2021 Globally-Robust Neural Networks
abstract
The threat of adversarial examples has motivated work on training certifiably robust neural networks to facilitate efficient verification of local robustness at inference time. We formalize a notion of global robustness, which captures the operational properties of on-line local robustness certification while yielding a natural learning objective for robust training. We show that widely-used architectures can be easily adapted to this objective by incorporating efficient global Lipschitz bounds into the network, yielding certifiably-robust models by construction that achieve state-of-the-art verifiable accuracy. Notably, this approach requires significantly less time and memory than recent certifiable training methods, and leads to negligible costs when certifying points on-line; for example, our evaluation shows that it is possible to train a large robust Tiny-Imagenet model in a matter of hours. Our models effectively leverage inexpensive global Lipschitz bounds for real-time certification, despite prior suggestions that tighter local bounds are needed for good performance; we posit this is possible because our models are specifically trained to achieve tighter global bounds. Namely, we prove that the maximum achievable verifiable accuracy for a given dataset is not improved by using a local bound.
Klas Leino, Zifan Wang 0001, Matt Fredrikson
ICML1
2021 Machine Learning Explainability and Robustness: Connected at the Hip
abstract
This tutorial examines the synergistic relationship between explainability methods for machine learning and a significant problem related to model quality: robustness against adversarial perturbations. We begin with a broad overview of approaches to explainable AI, before narrowing our focus to post-hoc explanation methods for predictive models. We discuss perspectives on what constitutes a "good'' explanation in various settings, with an emphasis on axiomatic justifications for various explanation methods. In doing so, we will highlight the importance of an explanation method's faithfulness to the target model, as this property allows one to distinguish between explanations that are unintelligible because of the method used to produce them, and cases where a seemingly poor explanation points to model quality issues. Next, we introduce concepts surrounding adversarial robustness, including adversarial attacks as well as a range of corresponding state-of-the-art defenses. Finally, building on the knowledge presented thus far, we present key insights from the recent literature on the connections between explainability and robustness, showing that many commonly-perceived explainability issues may be caused by non-robust model behavior. Accordingly, a careful study of adversarial examples and robustness can lead to models whose explanations better appeal to human intuition and domain knowledge.
Anupam Datta, Matt Fredrikson, Klas Leino, Kaiji Lu, Shayak Sen, Zifan Wang 0001
KDD3
2021 Relaxing Local Robustness
abstract
Certifiable local robustness, which rigorously precludes small-norm adversarial examples, has received significant attention as a means of addressing security concerns in deep learning. However, for some classification problems, local robustness is not a natural objective, even in the presence of adversaries; for example, if an image contains two classes of subjects, the correct label for the image may be considered arbitrary between the two, and thus enforcing strict separation between them is unnecessary. In this work, we introduce two relaxed safety properties for classifiers that address this observation: (1) relaxed top-k robustness, which serves as the analogue of top-k accuracy; and (2) affinity robustness, which specifies which sets of labels must be separated by a robustness margin, and which can be $\epsilon$-close in $\ell_p$ space. We show how to construct models that can be efficiently certified against each relaxed robustness property, and trained with very little overhead relative to standard gradient descent. Finally, we demonstrate experimentally that these relaxed variants of robustness are well-suited to several significant classification problems, leading to lower rejection rates and higher certified accuracies than can be obtained when certifying "standard" local robustness.
Klas Leino, Matt Fredrikson
NeurIPS1
2020 Influence Paths for Characterizing Subject-Verb Number Agreement in LSTM Language Models
abstract
LSTM-based recurrent neural networks are the state-of-the-art for many natural language processing (NLP) tasks.Despite their performance, it is unclear whether, or how, LSTMs learn structural features of natural languages such as subject-verb number agreement in English.Lacking this understanding, the generality of LSTMs on this task and their suitability for related tasks remains uncertain.Further, errors cannot be properly attributed to a lack of structural capability, training data omissions, or other exceptional faults.We introduce influence paths, a causal account of structural properties as carried by paths across gates and neurons of a recurrent neural network.The approach refines the notion of influence (the subject's grammatical number has influence on the grammatical number of the subsequent verb) into a set of gate-level or neuron-level paths.The set localizes and segments the concept (e.g., subject-verb agreement), its constituent elements (e.g., the subject), and related or interfering elements (e.g., attractors).We exemplify the methodology on a widely-studied multi-layer LSTM language model, demonstrating its accounting for subject-verb number agreement.The results offer both a finer and a more complete view of an LSTM's handling of this structural aspect of the English language than prior results based on diagnostic classifiers and ablation.
Kaiji Lu, Piotr Mardziel, Klas Leino, Matt Fredrikson, Anupam Datta
ACL3
2020 Stolen Memories: Leveraging Model Memorization for Calibrated White-Box Membership Inference
Klas Leino, Matt Fredrikson
USENIX Security Symposium1
2019 Feature-Wise Bias Amplification
Klas Leino, Emily Black, Matt Fredrikson, Shayak Sen, Anupam Datta
ICLR (Poster)1
2018 Influence-Directed Explanations for Deep Convolutional Networks
abstract
We study the problem of explaining a rich class of behavioral properties of deep neural networks. Distinctively, our influence-directed explanations approach this problem by peering inside the network to identify neurons with high influence on a quantity and distribution of interest, using an axiomatically-justified influence measure, and then providing an interpretation for the concepts these neurons represent. We evaluate our approach by demonstrating a number of its unique capabilities on convolutional neural networks trained on ImageNet. Our evaluation demonstrates that influence-directed explanations (1) identify influential concepts that generalize across instances, (2) can be used to extract the “essence” of what the network learned about a class, and (3) isolate individual features the network uses to make decisions and distinguish related classes.
Klas Leino, Shayak Sen, Anupam Datta, Matt Fredrikson
ITC1