Francesca Babiloni

dblp:169/4678 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
5since 2021 · last 2024
0000-0002-8447-6991ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Deep learning architectures and training · 39% 3D vision · 31% Generative modeling · 17%
Computer graphics and multimedia
2 papers
Image and video processing · 58% Visual content generation and editing · 42%

Topics — the 27 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
attention mechanism
1.632023
Linear Complexity Self-Attention With $3{\mathrm{rd}}$3 rd Order Polynomials · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Poly-NL: Linear Complexity Non-local Layers With 3rd Order Polynomials · ICCV 2021
TESA: Tensor Element Self-Attention via Matricization · CVPR 2020
Computer vision › 3D vision › 3d reconstruction › object reconstruction
3d head reconstruction
0.812024
ID-to-3D: Expressive ID-guided 3D Heads via Score Distillation Sampling · NeurIPS 2024
Machine learning › Generative modeling
diffusion model
0.812024
ID-to-3D: Expressive ID-guided 3D Heads via Score Distillation Sampling · NeurIPS 2024
Machine learning › Generative modeling › diffusion model
score distillation sampling
0.812024
ID-to-3D: Expressive ID-guided 3D Heads via Score Distillation Sampling · NeurIPS 2024
Visual content generation and editing
3d content creation
0.812024
ID-to-3D: Expressive ID-guided 3D Heads via Score Distillation Sampling · NeurIPS 2024
Computer vision › 3D vision › 3d face modeling
3d morphable model
0.712023
Adaptive Spiral Layers for Efficient 3D Representation Learning on Meshes · ICCV 2023
Computer vision › 3D vision
3d shape analysis
0.712023
Adaptive Spiral Layers for Efficient 3D Representation Learning on Meshes · ICCV 2023
Machine learning › Generative modeling › diffusion model
3d shape generation
0.712023
Adaptive Spiral Layers for Efficient 3D Representation Learning on Meshes · ICCV 2023
Machine learning › Deep learning architectures and training
convolutional neural network
0.712023
Tunable Convolutions with Parametric Multi-Loss Optimization · CVPR 2023
Computer vision › 3D vision
geometric deep learning
0.712023
Adaptive Spiral Layers for Efficient 3D Representation Learning on Meshes · ICCV 2023
Machine learning › Deep learning architectures and training › attention mechanism › efficient attention
linear complexity attention
0.712023
Linear Complexity Self-Attention With $3{\mathrm{rd}}$3 rd Order Polynomials · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computer vision › 3D vision › 3d shape representation › 3d shape representation learning
mesh representation learning
0.712023
Adaptive Spiral Layers for Efficient 3D Representation Learning on Meshes · ICCV 2023
Computer vision › 3D vision
shape matching
0.712023
Adaptive Spiral Layers for Efficient 3D Representation Learning on Meshes · ICCV 2023
Image and video processing
image restoration
0.712023
Tunable Convolutions with Parametric Multi-Loss Optimization · CVPR 2023
Image and video processing › image restoration
perception-distortion tradeoff
0.712023
Tunable Convolutions with Parametric Multi-Loss Optimization · CVPR 2023
Machine learning › Deep learning architectures and training › attention mechanism › self-attention
non-local block
0.512021
Poly-NL: Linear Complexity Non-local Layers With 3rd Order Polynomials · ICCV 2021
Computer vision › Image recognition and object detection
image classification
0.412020
TESA: Tensor Element Self-Attention via Matricization · CVPR 2020
Machine learning › Deep learning architectures and training › attention mechanism
self-attention
0.412020
TESA: Tensor Element Self-Attention via Matricization · CVPR 2020
Machine learning › Learning paradigms › continual learning
catastrophic forgetting
0.312018
Memory Aware Synapses: Learning What (not) to Forget · ECCV (3) 2018
Machine learning › Learning paradigms
continual learning
0.312018
Memory Aware Synapses: Learning What (not) to Forget · ECCV (3) 2018
Computer vision › Segmentation and scene understanding
instance segmentation
0.322021
Poly-NL: Linear Complexity Non-local Layers With 3rd Order Polynomials · ICCV 2021
TESA: Tensor Element Self-Attention via Matricization · CVPR 2020
Machine learning › Deep learning architectures and training › convolutional neural network
convolution design
0.212023
Adaptive Spiral Layers for Efficient 3D Representation Learning on Meshes · ICCV 2023
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.212023
Linear Complexity Self-Attention With $3{\mathrm{rd}}$3 rd Order Polynomials · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Visual content generation and editing
image-to-image translation
0.212023
Tunable Convolutions with Parametric Multi-Loss Optimization · CVPR 2023
Computer vision › Face, body and person analysis
face detection
0.112021
Poly-NL: Linear Complexity Non-local Layers With 3rd Order Polynomials · ICCV 2021
Computer vision › Image recognition and object detection
visual recognition
0.112021
Poly-NL: Linear Complexity Non-local Layers With 3rd Order Polynomials · ICCV 2021
Machine learning › Deep learning architectures and training › biologically plausible learning
synaptic plasticity
0.112018
Memory Aware Synapses: Learning What (not) to Forget · ECCV (3) 2018

Methods — techniques the papers use, named apart from their topics

score distillation sampling · 1.5neural parametric representation · 1.5diffusion model · 1.5parametric convolution · 1.3multi-loss optimization · 1.3polynomial approximation · 1.2summed-area tables · 0.7spiral convolution · 0.7polyadic decomposition · 0.7dynamic gating · 0.7
YearPublicationVenuePosition
2024 ID-to-3D: Expressive ID-guided 3D Heads via Score Distillation Sampling
abstract
We propose ID-to-3D, a method to generate identity- and text-guided 3D human heads with disentangled expressions, starting from even a single casually captured ‘in-the-wild’ image of a subject. The foundation of our approach is anchored in compositionality, alongside the use of task-specific 2D diffusion models as priors for optimization. First, we extend a foundational model with a lightweight expression-aware and ID-aware architecture, and create 2D priors for geometric and texture generation, via fine-tuning only 0.2% of its available training parameters. Then, we jointly leverage a neural parametric representation for the expression of each subject and a multi-stage generation of highly detailed geometry and albedo texture. This combination of strong face identity embeddings and our neural representation enables accurate reconstruction of not only facial features but also accessories and hair, and can be meshed to provide render-ready assets for gaming and telepresence. Our results achieve an unprecedented level of id-consistent and high-quality texture and geometry generation, generalizing to a ‘world’ of unseen 3D identities, without relying on large 3D captured datasets of human assets.
Francesca Babiloni, Alexander Lattas, Jiankang Deng, Stefanos Zafeiriou
NeurIPS1
2023 Tunable Convolutions with Parametric Multi-Loss Optimization
abstract
Behavior of neural networks is irremediably determined by the specific loss and data used during training. However it is often desirable to tune the model at inference time based on external factors such as preferences of the user or dynamic characteristics of the data. This is especially important to balance the perception-distortion trade-off of ill-posed image-to-image translation tasks. In this work, we propose to optimize a parametric tunable convolutional layer, which includes a number of different kernels, using a parametric multi-loss, which includes an equal number of objectives. Our key insight is to use a shared set of parameters to dynamically interpolate both the objectives and the kernels. During training, these parameters are sampled at random to explicitly optimize all possible combinations of objectives and consequently disentangle their effect into the corresponding kernels. During inference, these parameters become interactive inputs of the model hence enabling reliable and consistent control over the model behavior. Extensive experimental results demonstrate that our tunable convolutions effectively work as a drop-in replacement for traditional convolutions in existing neural networks at virtually no extra computational cost, outperforming state-of-the-art control strategies in a wide range of applications; including image denoising, deblurring, super-resolution, and style transfer.
Matteo Maggioni, Thomas Tanay, Francesca Babiloni, Steven McDonagh 0001, Ales Leonardis
CVPR3
2023 Adaptive Spiral Layers for Efficient 3D Representation Learning on Meshes
abstract
The success of deep learning models on structured data has generated significant interest in extending their application to non-Euclidean domains. In this work, we introduce a novel intrinsic operator suitable for representation learning on 3D meshes. Our operator is specifically tailored to adapt its behavior to the irregular structure of the underlying graph and effectively utilize its long-range dependencies, while at the same time ensuring computational efficiency and ease of optimization. In particular, inspired by the framework of Spiral Convolution, which extracts and transforms the vertices in the 3D mesh following a local spiral ordering, we propose a general operator that dynamically adjusts the length of the spiral trajectory and the parameters of the transformation for each processed vertex and mesh. Then, we use polyadic decomposition to factorize its dense weight tensor into a sequence of lighter linear layers that separately process features and vertices information, hence significantly reducing the computational complexity without introducing any stringent inductive biases. Notably, we leverage dynamic gating to achieve spatial adaptivity and induce global reasoning with constant time complexity benefitting from an efficient dynamic pooling mechanism based on Summed-Area-tables. Used as a drop-in replacement on existing architectures for shape correspondence our operator significantly improves the performance-efficiency trade-off, and in 3D shape generation with morphable models achieves state-of-the-art performance with a three-fold reduction in the number of parameters required. Project page: https://github.com/Fb2221/DFC
Francesca Babiloni, Matteo Maggioni, Thomas Tanay, Jiankang Deng, Ales Leonardis, Stefanos Zafeiriou
ICCV1
2023 Linear Complexity Self-Attention With $3{\mathrm{rd}}$3 rd Order Polynomials
abstract
Self-attention mechanisms and non-local blocks have become crucial building blocks for state-of-the-art neural architectures thanks to their unparalleled ability in capturing long-range dependencies in the input. However their cost is quadratic with the number of spatial positions hence making their use impractical in many real case applications. In this work, we analyze these methods through a polynomial lens, and we show that self-attention can be seen as a special case of a 3 rd order polynomial. Within this polynomial framework, we are able to design polynomial operators capable of accessing the same data pattern of non-local and self-attention blocks while reducing the complexity from quadratic to linear. As a result, we propose two modules (Poly-NL and Poly-SA) that can be used as "drop-in" replacements for more-complex non-local and self-attention layers in state-of-the-art CNNs and ViT architectures. Our modules can achieve comparable, if not better, performance across a wide range of computer vision tasks while keeping a complexity equivalent to a standard linear layer.
Francesca Babiloni, Ioannis Marras, Jiankang Deng, Filippos Kokkinos, Matteo Maggioni, Grigorios Chrysos 0002, Philip Torr 0001, Stefanos Zafeiriou
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Poly-NL: Linear Complexity Non-local Layers With 3rd Order Polynomials
abstract
Spatial self-attention layers, in the form of Non-Local blocks, introduce long-range dependencies in Convolutional Neural Networks by computing pairwise similarities among all possible positions. Such pairwise functions underpin the effectiveness of non-local layers, but also determine a complexity that scales quadratically with respect to the input size both in space and time. This is a severely limiting factor that practically hinders the applicability of non-local blocks to even moderately sized inputs. Previous works focused on reducing the complexity by modifying the underlying matrix operations, however in this work we aim to retain full expressiveness of non-local layers while keeping complexity linear. We overcome the efficiency limitation of non-local blocks by framing them as special cases of 3rd order polynomial functions. This fact enables us to formulate novel fast Non-Local blocks, capable of reducing the complexity from quadratic to linear with no loss in performance, by replacing any direct computation of pairwise similarities with element-wise multiplications. The proposed method, which we dub as "Poly-NL", is competitive with state-of-the-art performance across image recognition, instance segmentation, and face detection tasks, while having considerably less computational overhead.
Francesca Babiloni, Ioannis Marras, Filippos Kokkinos, Jiankang Deng, Grigorios Chrysos 0002, Stefanos Zafeiriou
ICCV1
2020 TESA: Tensor Element Self-Attention via Matricization
abstract
Representation learning is a fundamental part of modern computer vision, where abstract representations of data are encoded as tensors optimized to solve problems like image segmentation and inpainting. Recently, self-attention in the form of Non-Local Block has emerged as a powerful technique to enrich features, by capturing complex interdependencies in feature tensors. However, standard self-attention approaches leverage only spatial relationships, drawing similarities between vectors and overlooking correlations between channels. In this paper, we introduce a new method, called Tensor Element Self-Attention (TESA) that generalizes such work to capture interdependencies along all dimensions of the tensor using matricization. An order R tensor produces R results, one for each dimension. The results are then fused to produce an enriched output which encapsulates similarity among tensor elements. Additionally, we analyze self-attention mathematically, providing new perspectives on how it adjusts the singular values of the input feature tensor. With these new insights, we present experimental results demonstrating how TESA can benefit diverse problems including classification and instance segmentation. By simply adding a TESA module to existing networks, we substantially improve competitive baselines and set new state-of-the-art results for image inpainting on Celeb and low light raw-to-rgb image translation on SID.
Francesca Babiloni, Ioannis Marras, Gregory Slabaugh, Stefanos Zafeiriou
CVPR1
2018 Exploring the Challenges Towards Lifelong Fact Learning
Mohamed Elhoseiny 0001, Francesca Babiloni, Rahaf Aljundi, Marcus Rohrbach, Manohar Paluri, Tinne Tuytelaars
ACCV (6)2
2018 Memory Aware Synapses: Learning What (not) to Forget
Rahaf Aljundi, Francesca Babiloni, Mohamed Elhoseiny 0001, Marcus Rohrbach, Tinne Tuytelaars
ECCV (3)2
2017 Learning deep visual object models from noisy web data: How to make it work
abstract
Deep networks thrive when trained on large scale data collections. This has given ImageNet a central role in the development of deep architectures for visual object classification. However, ImageNet was created during a specific period in time, and as such it is prone to aging, as well as dataset bias issues. Moving beyond fixed training datasets will lead to more robust visual systems, especially when deployed on robots in new environments which must train on the objects they encounter there. To make this possible, it is important to break free from the need for manual annotators. Recent work has begun to investigate how to use the massive amount of images available on the Web in place of manual image annotations. We contribute to this research thread with two findings: (1) a study correlating a given level of noisily labels to the expected drop in accuracy, for two deep architectures, on two different types of noise, that clearly identifies GoogLeNet as a suitable architecture for learning from Web data; (2) a recipe for the creation of Web datasets with minimal noise and maximum visual variability, based on a visual and natural language processing concept expansion strategy. By combining these two results, we obtain a method for learning powerful deep object models automatically from the Web. We confirm the effectiveness of our approach through object categorization experiments using our Web-derived version of ImageNet on a popular robot vision benchmark database, and on a lifelong object discovery task on a mobile robot.
Nizar Massouh, Francesca Babiloni, Tatiana Tommasi, Jay Young, Nick Hawes, Barbara Caputo
IROS2