Brendan Duke

dblp:217/2928 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2023 Sparsifiner: Learning Sparse Instance-Dependent Attention for Efficient Vision Transformers
abstract
Vision Transformers (ViT) have shown competitive advantages in terms of performance compared to convolutional neural networks (CNNs), though they often come with high computational costs. To this end, previous methods explore different attention patterns by limiting a fixed number of spatially nearby tokens to accelerate the ViT's multi-head self-attention (MHSA) operations. However, such structured attention patterns limit the token-to-token connections to their spatial relevance, which disregards learned semantic connections from a full attention mask. In this work, we propose an approach to learn instance-dependent attention patterns, by devising a lightweight connectivity predictor module that estimates the connectivity score of each pair of tokens. Intuitively, two tokens have high connectivity scores if the features are considered relevant either spatially or semantically. As each token only attends to a small number of other tokens, the binarized connectivity masks are often very sparse by nature and therefore provide the opportunity to reduce network FLOPs via sparse computations. Equipped with the learned unstructured attention pattern, sparse attention ViT (Sparsifiner) produces a superior Pareto frontier between FLOPs and top-1 accuracy on ImageNet compared to token sparsity. Our method reduces 48% ~ 69% FLOPs of MHSA while the accuracy drop is within 0.4%. We also show that combining attention and token sparsity reduces ViT FLOPs by over 60%.
Cong Wei 0001, Brendan Duke, Ruowei Jiang, Parham Aarabi, Graham W. Taylor, Florian Shkurti
CVPR2
2022 Exploring Gradient-Based Multi-directional Controls in GANs
Ruowei Jiang, Brendan Duke, Han Zhao 0002, Parham Aarabi
ECCV (23)3
2022 Synthesizing ultraviolet skin images via GAN with Gaussian weighted patch blending
abstract
In this work, we explore a novel application of synthesizing ultraviolet skin images from RGB images using an unpaired training framework for image-to-image translation. To synthesize high resolution outputs, we propose a novel Gaussian-based patch blending technique that is designed following the characteristics of GANs. Specifically, we weigh the pixels at the same coordinates among multiple generated patches based on their distance to the center point. Our proposed method is performant, taking 0.72s for a whole face image at resolution 960x720 at inference time and generates realistic-looking ultraviolet images. We also show high correspondence of our synthesized images with the true ultraviolet images qualitatively. Finally, our novel blending approach achieves significant improvements compared with other blending methods.
Ruowei Jiang, Brendan Duke, Frédéric Flament, Parham Aarabi
ISM2
2022 Real-time Virtual-Try-On from a Single Example Image through Deep Inverse Graphics and Learned Differentiable Renderers
abstract
Abstract Augmented reality applications have rapidly spread across online retail platforms and social media, allowing consumers to virtually try‐on a large variety of products, such as makeup, hair dying, or shoes. However, parametrizing a renderer to synthesize realistic images of a given product remains a challenging task that requires expert knowledge. While recent work has introduced neural rendering methods for virtual try‐on from example images, current approaches are based on large generative models that cannot be used in real‐time on mobile devices. This calls for a hybrid method that combines the advantages of computer graphics and neural rendering approaches. In this paper, we propose a novel framework based on deep learning to build a real‐time inverse graphics encoder that learns to map a single example image into the parameter space of a given augmented reality rendering engine. Our method leverages self‐supervised learning and does not require labeled training data, which makes it extendable to many virtual try‐on applications. Furthermore, most augmented reality renderers are not differentiable in practice due to algorithmic choices or implementation constraints to reach real‐time on portable devices. To relax the need for a graphics‐based differentiable renderer in inverse graphics problems, we introduce a trainable imitator module. Our imitator is a generative network that learns to accurately reproduce the behavior of a given non‐differentiable renderer. We propose a novel rendering sensitivity loss to train the imitator, which ensures that the network learns an accurate and continuous representation for each rendering parameter. Automatically learning a differentiable renderer, as proposed here, could be beneficial for various inverse graphics tasks. Our framework enables novel applications where consumers can virtually try‐on a novel unknown product from an inspirational reference image on social media. It can also be used by computer graphics artists to automatically create realistic rendering from a reference product image.
Robin Kips, Ruowei Jiang, Sileye O. Ba, Brendan Duke, Matthieu Perrot, Pietro Gori, Isabelle Bloch
Comput. Graph. Forum4
2021 SSTVOS: Sparse Spatiotemporal Transformers for Video Object Segmentation
abstract
In this paper we introduce a Transformer-based approach to video object segmentation (VOS). To address compounding error and scalability issues of prior work, we propose a scalable, end-to-end method for VOS called Sparse Spatiotemporal Transformers (SST). SST extracts per-pixel representations for each object in a video using sparse attention over spatiotemporal features. Our attention-based formulation for VOS allows a model to learn to attend over a history of multiple frames and provides suitable inductive bias for performing correspondence-like computations necessary for solving motion segmentation. We demonstrate the effectiveness of attention-based over recurrent networks in the spatiotemporal domain. Our method achieves competitive results on YouTube-VOS and DAVIS 2017 with improved scalability and robustness to occlusions compared with the state of the art. Code is available at https://github.com/dukebw/SSTVOS.
Brendan Duke, Abdalla Ahmed, Christian Wolf 0001, Parham Aarabi, Graham W. Taylor
CVPR1
2021 LOHO: Latent Optimization of Hairstyles via Orthogonalization
abstract
Hairstyle transfer is challenging due to hair structure differences in the source and target hair. Therefore, we propose Latent Optimization of Hairstyles via Orthogonalization (LOHO), an optimization-based approach using GAN inversion to infill missing hair structure details in latent space during hairstyle transfer. Our approach decomposes hair into three attributes: perceptual structure, appearance, and style, and includes tailored losses to model each of these attributes independently. Furthermore, we propose two-stage optimization and gradient orthogonalization to enable disentangled latent space optimization of our hair attributes. Using LOHO for latent space manipulation, users can synthesize novel photorealistic images by manipulating hair attributes either individually or jointly, transferring the desired attributes from reference hairstyles. LOHO achieves a superior FID compared with the current state-of-the-art (SOTA) for hairstyle transfer. Additionally, LOHO preserves the subject’s identity comparably well according to PSNR and SSIM when compared to SOTA image embedding pipelines. Code is available at https://github.com/dukebw/LOHO.
Rohit Saha, Brendan Duke, Florian Shkurti, Graham W. Taylor, Parham Aarabi
CVPR2