Michela Paganini

dblp:210/2609 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0003-4102-8002ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 19% Trustworthy machine learning · 16% Deep learning architectures and training · 15%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › alignment
pluralistic alignment
0.912025
Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models · NeurIPS 2025
Machine learning › Trustworthy machine learning
safety evaluation
0.912025
Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.912025
Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models · NeurIPS 2025
Machine learning › Representation and self-supervised learning › representation learning
invariant representation learning
0.712023
Neural Algorithmic Reasoning with Causal Regularisation · ICML 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge incorporation › knowledge-infused learning › neuro-symbolic learning
neural algorithmic reasoning
0.712023
Neural Algorithmic Reasoning with Causal Regularisation · ICML 2023
Natural language and speech › Question answering and dialogue systems
knowledge-intensive tasks
0.612022
Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022
Machine learning › Deep learning architectures and training
mixture of experts
0.612022
Unified Scaling Laws for Routed Language Models · ICML 2022
Natural language and speech › Language models and text generation
retrieval-augmented language models
0.612022
Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022
Machine learning › Deep learning architectures and training
scaling laws
0.612022
Unified Scaling Laws for Routed Language Models · ICML 2022
Information retrieval
document retrieval
0.612022
Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022
Machine learning › Efficient and distributed learning › model compression › sparse training
lottery ticket hypothesis
0.412019
One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers · NeurIPS 2019
Machine learning › Efficient and distributed learning
model compression
0.412019
One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers · NeurIPS 2019
Machine learning › Trustworthy machine learning
out-of-distribution generalization
0.212023
Neural Algorithmic Reasoning with Causal Regularisation · ICML 2023
Machine learning › Trustworthy machine learning
robustness
0.212023
Neural Algorithmic Reasoning with Causal Regularisation · ICML 2023

Methods — techniques the papers use, named apart from their topics

differentiable encoder · 1.1chunked cross-attention · 1.1human evaluation · 0.9LLM judgment · 0.9self-supervised learning · 0.7data augmentation · 0.7causal regularization · 0.7power-law scaling · 0.6effective parameter count · 0.6pruning · 0.4
YearPublicationVenuePosition
2025 Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models
abstract
Current text-to-image (T2I) models often fail to account for diverse human experiences, leading to misaligned systems. We advocate for pluralism in AI alignment, where an AI understands and is steerable towards diverse, and often conflicting, human values. Our work provides three core contributions to achieve this in T2I models. First, we introduce a novel dataset for Diverse Intersectional Visual Evaluation (DIVE) -- the first multimodal dataset for pluralistic alignment. It enables deep alignment to diverse safety perspectives through a large pool of demographically intersectional human raters who provided extensive feedback across 1000 prompts, with high replication, capturing nuanced safety perceptions. Second, we empirically confirm demographics as a crucial proxy for diverse viewpoints in this domain, revealing significant, context-dependent differences in harm perception that diverge from conventional evaluations. Finally, we discuss implications for building aligned T2I models, including efficient data collection strategies, LLM judgment capabilities, and model steerability towards diverse perspectives. This research offers foundational tools for more equitable and aligned T2I systems.Content Warning: The paper includes sensitive content that may be harmful.
Charvi Rastogi, Tian Huey Teh, Pushkar Mishra, Roma Patel, Ding Wang 0006, Mark Diaz, Alicia Parrish, Aida Mostafazadeh Davani, Zoe Ashwood, Michela Paganini, Vinodkumar Prabhakaran, Verena Rieser, Lora Aroyo
NeurIPS10
2023 Neural Algorithmic Reasoning with Causal Regularisation
abstract
Recent work on neural algorithmic reasoning has investigated the reasoning capabilities of neural networks, effectively demonstrating they can learn to execute classical algorithms on unseen data coming from the train distribution. However, the performance of existing neural reasoners significantly degrades on out-of-distribution (OOD) test data, where inputs have larger sizes. In this work, we make an important observation: there are many different inputs for which an algorithm will perform certain intermediate computations identically. This insight allows us to develop data augmentation procedures that, given an algorithm’s intermediate trajectory, produce inputs for which the target algorithm would have exactly the same next trajectory step. We ensure invariance in the next-step prediction across such inputs, by employing a self-supervised objective derived by our observation, formalised in a causal graph. We prove that the resulting method, which we call Hint-ReLIC, improves the OOD generalisation capabilities of the reasoner. We evaluate our method on the CLRS algorithmic reasoning benchmark, where we show up to 3x improvements on the OOD test data.
Beatrice Bevilacqua, Kyriacos Nikiforou, Borja Ibarz, Ioana Bica, Michela Paganini, Charles Blundell, Jovana Mitrovic, Petar Velickovic
ICML5
2022 Improving Language Models by Retrieving from Trillions of Tokens
abstract
We enhance auto-regressive language models by conditioning on document chunks retrieved from a large corpus, based on local similarity with preceding tokens. With a 2 trillion token database, our Retrieval-Enhanced Transformer (RETRO) obtains comparable performance to GPT-3 and Jurassic-1 on the Pile, despite using 25{\texttimes} fewer parameters. After fine-tuning, RETRO performance translates to downstream knowledge-intensive tasks such as question answering. RETRO combines a frozen Bert retriever, a differentiable encoder and a chunked cross-attention mechanism to predict tokens based on an order of magnitude more data than what is typically consumed during training. We typically train RETRO from scratch, yet can also rapidly RETROfit pre-trained transformers with retrieval and still achieve good performance. Our work opens up new avenues for improving language models through explicit memory at unprecedented scale.
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche 0002, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Albin Cassirer, Andrew Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, Laurent Sifre
ICML21
2022 Unified Scaling Laws for Routed Language Models
abstract
The performance of a language model has been shown to be effectively modeled as a power-law in its parameter count. Here we study the scaling behaviors of Routing Networks: architectures that conditionally use only a subset of their parameters while processing an input. For these models, parameter count and computational requirement form two independent axes along which an increase leads to better performance. In this work we derive and justify scaling laws defined on these two variables which generalize those known for standard language models and describe the performance of a wide range of routing architectures trained via three different techniques. Afterwards we provide two applications of these laws: first deriving an Effective Parameter Count along which all models scale at the same rate, and then using the scaling coefficients to give a quantitative comparison of the three routing techniques considered. Our analysis derives from an extensive evaluation of Routing Networks across five orders of magnitude of size, including models with hundreds of experts and hundreds of billions of parameters.
Aidan Clark, Diego de Las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake A. Hechtman, Trevor Cai, Sebastian Borgeaud, George van den Driessche 0002, Eliza Rutherford, Tom Hennigan, Matthew J. Johnson 0002, Albin Cassirer, Elena Buchatskaya, David Budden, Laurent Sifre, Simon Osindero, Oriol Vinyals, Marc'Aurelio Ranzato, Jack W. Rae, Erich Elsen, Koray Kavukcuoglu, Karen Simonyan
ICML5
2019 One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers
abstract
The success of lottery ticket initializations (Frankle and Carbin, 2019) suggests that small, sparsified networks can be trained so long as the network is initialized appropriately. Unfortunately, finding these "winning ticket'' initializations is computationally expensive. One potential solution is to reuse the same winning tickets across a variety of datasets and optimizers. However, the generality of winning ticket initializations remains unclear. Here, we attempt to answer this question by generating winning tickets for one training configuration (optimizer and dataset) and evaluating their performance on another configuration. Perhaps surprisingly, we found that, within the natural images domain, winning ticket initializations generalized across a variety of datasets, including Fashion MNIST, SVHN, CIFAR-10/100, ImageNet, and Places365, often achieving performance close to that of winning tickets generated on the same dataset. Moreover, winning tickets generated using larger datasets consistently transferred better than those generated using smaller datasets. We also found that winning ticket initializations generalize across optimizers with high performance. These results suggest that winning ticket initializations generated by sufficiently large datasets contain inductive biases generic to neural networks more broadly which improve training across many settings and provide hope for the development of better initialization methods.
Ari S. Morcos, Haonan Yu, Michela Paganini, Yuandong Tian
NeurIPS3