Andrea Schioppa

dblp:228/8404 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Trustworthy machine learning · 61% Optimization for machine learning · 13% Machine translation · 9%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
1.422024
Efficient Sketches for Training Data Attribution and Studying the Loss Landscape · NeurIPS 2024
Theoretical and Practical Perspectives on what Influence Functions Do · NeurIPS 2023
Machine learning › Optimization for machine learning
sketching
0.812024
Efficient Sketches for Training Data Attribution and Studying the Loss Landscape · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability
training data attribution
0.812024
Efficient Sketches for Training Data Attribution and Studying the Loss Landscape · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability › training data attribution
influence function
0.712023
Theoretical and Practical Perspectives on what Influence Functions Do · NeurIPS 2023
Machine learning › Trustworthy machine learning › interpretability
model debugging
0.712023
Theoretical and Practical Perspectives on what Influence Functions Do · NeurIPS 2023
Machine learning › Generative modeling › diffusion model › controllable generation
attribute-controlled generation
0.512021
Controlling Machine Translation for Multiple Attributes with Additive Interventions · EMNLP (1) 2021
Natural language and speech › Machine translation
controllable machine translation
0.512021
Controlling Machine Translation for Multiple Attributes with Additive Interventions · EMNLP (1) 2021
Natural language and speech › Language models and text generation
text generation
0.512021
Controlling Machine Translation for Multiple Attributes with Additive Interventions · EMNLP (1) 2021

Methods — techniques the papers use, named apart from their topics

fine-tuning · 1.2intrinsic dimension estimation · 0.8hessian vector product sketching · 0.8influence functions · 0.7vector-valued intervention · 0.5
YearPublicationVenuePosition
2024 Efficient Sketches for Training Data Attribution and Studying the Loss Landscape
abstract
The study of modern machine learning models often necessitates storing vast quantities of gradients or Hessian vector products (HVPs). Traditional sketching methods struggle to scale under these memory constraints. We present a novel framework for scalable gradient and HVP sketching, tailored for modern hardware. We provide theoretical guarantees and demonstrate the power of our methods in applications like training data attribution, Hessian spectrum analysis, and intrinsic dimension computation for pre-trained language models. Our work sheds new light on the behavior of pre-trained language models, challenging assumptions about their intrinsic dimensionality and Hessian properties.
Andrea Schioppa
NeurIPS1
2023 Theoretical and Practical Perspectives on what Influence Functions Do
abstract
Influence functions (IF) have been seen as a technique for explaining model predictions through the lens of the training data. Their utility is assumed to be in identifying training examples "responsible" for a prediction so that, for example, correcting a prediction is possible by intervening on those examples (removing or editing them) and retraining the model. However, recent empirical studies have shown that the existing methods of estimating IF predict the leave-one-out-and-retrain effect poorly. In order to understand the mismatch between the theoretical promise and the practical results, we analyse five assumptions made by IF methods which are problematic for modern-scale deep neural networks and which concern convexity, numeric stability, training trajectory and parameter divergence. This allows us to clarify what can be expected theoretically from IF. We show that while most assumptions can be addressed successfully, the parameter divergence poses a clear limitation on the predictive power of IF: influence fades over training time even with deterministic training. We illustrate this theoretical result with BERT and ResNet models. Another conclusion from the theoretical analysis is that IF are still useful for model debugging and correcting even though some of the assumptions made in prior work do not hold: using natural language processing and computer vision tasks, we verify that mis-predictions can be successfully corrected by taking only a few fine-tuning steps on influential examples.
Andrea Schioppa, Katja Filippova, Ivan Titov 0001, Polina Zablotskaia
NeurIPS1
2022 Scaling Up Influence Functions
abstract
We address efficient calculation of influence functions for tracking predictions back to the training data. We propose and analyze a new approach to speeding up the inverse Hessian calculation based on Arnoldi iteration. With this improvement, we achieve, to the best of our knowledge, the first successful implementation of influence functions that scales to full-size (language and vision) Transformer models with several hundreds of millions of parameters. We evaluate our approach in image classification and sequence-to-sequence tasks with tens to a hundred of millions of training examples. Our code is available at https://github.com/google-research/jax-influence.
Andrea Schioppa, Polina Zablotskaia, David Vilar, Artem Sokolov 0001
AAAI1
2021 Controlling Machine Translation for Multiple Attributes with Additive Interventions
abstract
Fine-grained control of machine translation (MT) outputs along multiple attributes is critical for many modern MT applications and is a requirement for gaining users' trust.A standard approach for exerting control in MT is to prepend the input with a special tag to signal the desired output attribute.Despite its simplicity, attribute tagging has several drawbacks: continuous values must be binned into discrete categories, which is unnatural for certain applications; interference between multiple tags is poorly understood.We address these problems by introducing vector-valued interventions which allow for fine-grained control over multiple attributes simultaneously via a weighted linear combination of the corresponding vectors.For some attributes, our approach even allows for fine-tuning a model trained without annotations to support such interventions.In experiments with three attributes (length, politeness and monotonicity) and two language pairs (English to German and Japanese) our models achieve better control over a wider range of tasks compared to tagging, and translation quality does not degrade when no control is requested.Finally, we demonstrate how to enable control in an already trained model after a relatively cheap fine-tuning stage.* Google AI Resident.
Andrea Schioppa, David Vilar, Artem Sokolov 0001, Katja Filippova
EMNLP (1)1