Fabio Pizzati

dblp:241/5366 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0002-6249-5649ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 9 since 2021
YearPublicationVenuePosition
2025 Video Motion Transfer with Diffusion Transformers
abstract
We propose DiTFlow, a method for transferring the motion of a reference video to a newly synthesized one, designed specifically for Diffusion Transformers (DiT). We first process the reference video with a pre-trained DiT to analyze cross-frame attention maps and extract a patch-wise motion signal called the Attention Motion Flow (AMF). We guide the latent denoising process in an optimization-based, training-free, manner by optimizing latents with our AMF loss to generate videos reproducing the motion of the reference one. We also apply our optimization strategy to transformer positional embeddings, granting us a boost in zero-shot motion transfer capabilities. We evaluate DiTFlow against recently published methods, outperforming all across multiple metrics and human evaluation.
Alexander Pondaven, Aliaksandr Siarohin, Sergey Tulyakov, Philip Torr 0001, Fabio Pizzati
CVPR5
2025 AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
Runtao Liu, Chen I Chieh, Jindong Gu, Renjie Pi, Qifeng Chen 0001, Philip Torr 0001, Ashkan Khakzar, Fabio Pizzati
ICCV9
2025 MatchDiffusion: Training-Free Generation of Match-Cuts
abstract
Match-cuts are powerful cinematic tools that create seamless transitions between scenes, delivering strong visual and metaphorical connections. However, crafting match-cuts is a challenging, resource-intensive process requiring deliberate artistic planning. In MatchDiffusion, we present the first training-free method for match-cut generation using text-to-video diffusion models. MatchDiffusion leverages a key property of diffusion models: early denoising steps define the scene's broad structure, while later steps add details. Guided by this insight, MatchDiffusion employs "Joint Diffusion" to initialize generation for two prompts from shared noise, aligning structure and motion. It then applies "Disjoint Diffusion", allowing the videos to diverge and introduce unique details. This approach produces visually coherent videos suited for match-cuts. User studies and metrics demonstrate MatchDiffusion's effectiveness and potential to democratize match-cut creation.
Alejandro Pardo, Fabio Pizzati, Tong Zhang 0001, Alexander Pondaven, Philip Torr 0001, Juan C. Pérez, Bernard Ghanem
ICCV2
2025 LaCoOT: Layer Collapse through Optimal Transport
abstract
Although deep neural networks are well-known for their outstanding performance in tackling complex tasks, their hunger for computational resources remains a significant hurdle, posing energy-consumption issues and restricting their deployment on resource-constrained devices, preventing their widespread adoption. In this paper, we present an optimal transport-based method to reduce the depth of over-parametrized deep neural networks, alleviating their computational burden. More specifically, we propose a new regularization strategy based on the Max-Sliced Wasserstein distance to minimize the distance between the intermediate feature distributions in the neural network. We show that minimizing this distance enables the complete removal of intermediate layers in the network, achieving better performance/depth trade-off compared to existing techniques. We assess the effectiveness of our method on traditional image classification setups and extend it to generative image models. Our code is available at https://github.com/VGCQ/LaCoOT.
Victor Quétu, Zhu Liao, Nour Hezbri, Fabio Pizzati, Enzo Tartaglione
ICCV4
2025 Towards Reliable Identification of Diffusion-based Image Manipulations
abstract
Changing facial expressions, gestures, or background details may dramatically alter the meaning conveyed by an image. Notably, recent advances in diffusion models greatly improve the quality of image manipulation while also opening the door to misuse. Identifying changes made to authentic images, thus, becomes an important task, constantly challenged by new diffusion-based editing tools. To this end, we propose a novel approach for ReliAble iDentification of inpainted AReas (RADAR). RADAR builds on existing foundation models and combines features from different image modalities. It also incorporates an auxiliary contrastive loss that helps to isolate manipulated image patches. We demonstrate these techniques to significantly improve both the accuracy of our method and its generalisation to a large number of diffusion models. To support realistic evaluation, we further introduce BBC-PAIR, a new comprehensive benchmark, with images tampered by 28 diffusion models. Our experiments show that RADAR achieves excellent results, outperforming the state-of-the-art in detecting and localising image edits made by both seen and unseen diffusion models. Further information about our code, data and models, including separate licensing terms, will be publicly available at https://alex-costanzino.github.io/radar/.
Alex Costanzino, Woody Bayliss, Juil Sock, Marc Gorriz, Danijela Horak, Ivan Laptev, Philip Torr 0001, Fabio Pizzati
NeurIPS8
2024 Material Palette: Extraction of Materials from a Single Image
abstract
Physically-Based Rendering (PBR) is key to modeling the interaction between light and materials, and finds extensive applications across computer graphics domains. However, acquiring PBR materials is costly and requires special apparatus. In this paper, we propose a method to extract PBR materials from a single real-world image. We do so in two steps: first, we map regions of the image to material concept tokens using a diffusion model, allowing the sampling of texture images resembling each material in the scene. Second, we leverage a separate network to decom-pose the generated textures into spatially varying BRDFs (SVBRDFs), offering us readily usable materials for rendering applications. Our approach relies on existing synthetic material libraries with SVBRDF ground truth. It exploits a diffusion-generated RGB texture dataset to allow generalization to new samples using unsupervised do-main adaptation (UDA). Our contributions are thoroughly evaluated on synthetic and real-world datasets. We further demonstrate the applicability of our method for editing 3D scenes with materials estimated from real photographs. Along with video, we share code and models as open-source on the project page: https://github.com/astra-vision/MaterialPalette.
Ivan Lopes, Fabio Pizzati, Raoul de Charette
CVPR2
2024 On Pretraining Data Diversity for Self-Supervised Learning
Hasan Hammoud, Tuhin Das, Fabio Pizzati, Philip Torr 0001, Adel Bibi, Bernard Ghanem
ECCV (56)3
2024 Latent Guard: A Safety Framework for Text-to-Image Generation
Runtao Liu, Ashkan Khakzar, Jindong Gu, Qifeng Chen 0001, Philip Torr 0001, Fabio Pizzati
ECCV (26)6
2024 Position: Near to Mid-term Risks and Opportunities of Open-Source Generative AI
abstract
In the next few years, applications of Generative AI are expected to revolutionize a number of different areas, ranging from science & medicine to education. The potential for these seismic changes has triggered a lively debate about potential risks and resulted in calls for tighter regulation, in particular from some of the major tech companies who are leading in AI development. While regulation is important, it is key that it does not put at risk the budding field of open-source Generative AI. We argue for the responsible open sourcing of generative AI models in the near and medium term. To set the stage, we first introduce an AI openness taxonomy system and apply it to 40 current large language models. We then outline differential benefits and risks of open versus closed source AI and present potential risk mitigation, ranging from best practices to calls for technical and scientific contributions. We hope that this report will add a much needed missing voice to the current public discourse on near to mid-term AI safety and other societal impact.
Francisco Girbal Eiras, Aleksandar Petrov, Bertie Vidgen, Christian Schröder de Witt, Fabio Pizzati, Katherine Elkins, Supratik Mukhopadhyay, Adel Bibi, Botos Csaba, Fabro Steibel, Fazl Barez, Genevieve Smith, Gianluca Guadagni, Jon Chun, Jordi Cabot, Joseph Marvin Imperial, Juan A. Nolazco-Flores, Lori Landay, Matthew Thomas Jackson, Paul Röttger, Philip Torr 0001, Trevor Darrell, Jakob N. Foerster
ICML5
2023 Physics-Informed Guided Disentanglement in Generative Networks
abstract
Image-to-image translation (i2i) networks suffer from entanglement effects in presence of physics-related phenomena in target domain (such as occlusions, fog, etc), lowering altogether the translation quality, controllability and variability. In this paper, we propose a general framework to disentangle visual traits in target images. Primarily, we build upon collection of simple physics models, guiding the disentanglement with a physical model that renders some of the target traits, and learning the remaining ones. Because physics allows explicit and interpretable outputs, our physical models (optimally regressed on target) allows generating unseen scenarios in a controllable manner. Secondarily, we show the versatility of our framework to neural-guided disentanglement where a generative network is used in place of a physical model in case the latter is not directly accessible. Altogether, we introduce three strategies of disentanglement being guided from either a fully differentiable physics model, a (partially) non-differentiable physics model, or a neural network. The results show our disentanglement strategies dramatically increase performances qualitatively and quantitatively in several challenging scenarios for image translation.
Fabio Pizzati, Pietro Cerri, Raoul de Charette
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 ManiFest: Manifold Deformation for Few-Shot Image Translation
Fabio Pizzati, Jean-François Lalonde, Raoul de Charette
ECCV (17)1
2021 CoMoGAN: Continuous Model-Guided Image-to-Image Translation
abstract
CoMoGAN is a continuous GAN relying on the unsupervised reorganization of the target data on a functional manifold. To that matter, we introduce a new Functional Instance Normalization layer and residual mechanism, which together disentangle image content from position on target manifold. We rely on naive physics-inspired models to guide the training while allowing private model/translations features. CoMoGAN can be used with any GAN backbone and allows new types of image translation, such as cyclic image translation like timelapse generation, or detached linear translation. On all datasets, it outperforms the literature. Our code is available in this page: https://github.com/cv-rits/CoMoGAN.
Fabio Pizzati, Pietro Cerri, Raoul de Charette
CVPR1
2020 Model-Based Occlusion Disentanglement for Image-to-Image Translation
Fabio Pizzati, Pietro Cerri, Raoul de Charette
ECCV (20)1
2020 Domain Bridge for Unpaired Image-to-Image Translation and Unsupervised Domain Adaptation
abstract
Image-to-image translation architectures may have limited effectiveness in some circumstances. For example, while generating rainy scenarios, they may fail to model typical traits of rain as water drops, and this ultimately impacts the synthetic images realism. With our method, called domain bridge, web-crawled data are exploited to reduce the domain gap, leading to the inclusion of previously ignored elements in the generated images. We make use of a network for clear to rain translation trained with the domain bridge to extend our work to Unsupervised Domain Adaptation (UDA). In that context, we introduce an online multimodal style-sampling strategy, where image translation multimodality is exploited at training time to improve performances. Finally, a novel approach for self-supervised learning is presented, and used to further align the domains. With our contributions, we simultaneously increase the realism of the generated images, while reaching on par performances with respect to the UDA state-of-the-art, with a simpler approach.
Fabio Pizzati, Raoul de Charette, Michela Zaccaria, Pietro Cerri
WACV1
2019 Enhanced free space detection in multiple lanes based on single CNN with scene identification
abstract
Many systems for autonomous vehicles' navigation rely on lane detection. Traditional algorithms usually estimate only the position of the lanes on the road, but an autonomous control system may also need to know if a lane marking can be crossed or not, and what portion of space inside the lane is free from obstacles, to make safer control decisions. On the other hand, free space detection algorithms only detect navigable areas, without information about lanes. State-of-the-art algorithms use CNNs for both tasks, with significant consumption of computing resources. We propose a novel approach that estimates the free space inside each lane, with a single CNN. Additionally, adding only a small requirement concerning GPU RAM, we infer the road type, that will be useful for path planning. To achieve this result, we train a multi-task CNN. Then, we further elaborate the output of the network, to extract polygons that can be effectively used in navigation control. Finally, we provide a computationally efficient implementation, based on ROS, that can be executed in real time. Our code and trained models are available online.
Fabio Pizzati
IV1
2019 A Deep-Learning Approach for Parking Slot Detection on Surround-View Images
abstract
Being able to automatically detect parking slots during navigation is an important part of an autonomous driving system. Most of the current parking slot detectors either work on images coming from static cameras or are based on hand-designed low-level visual features, limiting the domains of application of such systems considerably. Learning the most suitable features for the task directly from data would allow the system to function in a broader range of situations and be more robust to noise and different observation conditions. This paper presents an end-to-end deep neural network trained to perform automatic parking slot detection and classification on surround-view images, generated by fusing the views of four different cameras positioned on a vehicle. The network architecture is based on the Faster R-CNN baseline. To account for the fact that parking slots can be of different shapes and can be observed at different angles, the predicted bounding boxes are generic quadrilaterals rather than image-aligned rectangles. Moreover, instead of computing offsets wrt anchor boxes, the RPN regresses the region proposals directly. In order to train the network, a small training set comprised of a few hundred images has been manually annotated. The network shows good performance and generalization capability on new examples containing parking slots of the same kinds as those in the training set, suggesting that a wider array of parking slot types could easily be integrated into the system by expanding the dataset.
Andrea Zinelli, Luigi Musto, Fabio Pizzati
IV3