EDBT 2026 Demo / reviewers in the wild / expert
Ruowei Jiang
dblp:210/7023
· DBLP profile ↗
9ranked-venue papers
1as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards High-Fidelity, Identity-Preserving Real-Time Makeup Transfer: Decoupling Style GenerationabstractWe propose a framework for real-time virtual makeup transfer that achieves high-fidelity, identity-preserving results with strong temporal consistency. Existing methods often struggle to disentangle semi-transparent makeup from skin tones and other identity features, leading to identity shifts and fairness concerns. Furthermore, they also lack real-time capabilities and fail to maintain temporal consistency, limiting adoption in practical virtual try-on applications. To address these challenges, we decouple makeup transfer into two stages: transparent makeup mask extraction and graphics-based real-time makeup rendering. Once extracted, makeup masks can be applied in real-time for live video try-on. We generate pseudo-ground-truth data via a hybrid graphics-based rendering pipeline and an unsupervised clustering method, enabling robust training without real paired before-and-after makeup data. To further enhance transparency estimation and color fidelity, we propose transparency-aware reconstruction and lip color objectives. Our method consistently transfers fine-grained makeup details across diverse skin tones and expressions while maintaining temporal stability. Experiments demonstrate superior accuracy, stability, and efficiency over state-of-the-art baselines, making our approach practical for live virtual try-on applications. Video demonstrations are available in supplementary material. Kin Ching Lydia Chau, Ruowei Jiang |
WACV | 3 |
| 2024 | SCE-MAE: Selective Correspondence Enhancement with Masked Autoencoder for Self-Supervised Landmark EstimationabstractSelf-supervised landmark estimation is a challenging task that demands the formation of locally distinct feature representations to identify sparse facial landmarks in the absence of annotated data. To tackle this task, existing state-of-the-art (SOTA) methods (1) extract coarse features from backbones that are trained with instance-level self-supervised learning (SSL) paradigms, which neglect the dense prediction nature of the task, (2) aggregate them into memory-intensive hypercolumn formations, and (3) su-pervise lightweight projector networks to naï vely establish full local correspondences among all pairs of spatial features. In this paper, we introduce SCE-MAE, a framework that (1) leverages the MAE [14], a region-level SSL method that naturally better suits the landmark prediction task, (2) operates on the vanilla feature map instead of on expen-sive hypercolumns, and (3) employs a Correspondence Ap-proximation and Refinement Block (CARB) that utilizes a simple density peak clustering algorithm and our proposed Locality-Constrained Repellence Loss to directly hone only select local correspondences. We demonstrate through extensive experiments that SCE-MAE is highly effective and robust, outperforming existing SOTA methods by large mar-gins of ~20%-44% on the landmark matching and ~9%-15% on the landmark detection tasks. Kejia Yin, Varshanth S. Rao, Ruowei Jiang, Parham Aarabi, David B. Lindell |
CVPR | 3 |
| 2024 | Occlusion-Aware Real-Time Tiny Facial Alignment Model for Makeup Virtual Try-OnabstractReal-time makeup virtual try-on (VTO) on resource-constrained platforms like mobile devices and web browsers demands a delicate balance: models must be accurate enough for realistic results yet lightweight and fast enough for smooth performance. Existing approaches often rely on separate models for facial landmark detection and occlusion-aware segmentation, increasing complexity and hindering real-time performance. To address this, we propose a novel, unified model that performs both tasks within a single, highly efficient architecture. Specifically designed for VTO, our model offers enhanced accuracy around critical areas like the eyes and lips. We further optimize for real-time performance by leveraging temporal information: predictions from previous video frames guide current predictions, increasing parallelism and reducing inference time to as little as 16ms on an iPhone 14. Trained with a simplified pipeline, our unified model achieves accuracy comparable to state-of-the-art lightweight alignment models while maintaining a small footprint. Kin Ching Lydia Chau, Ruowei Jiang |
ISM | 3 |
| 2023 | Sparsifiner: Learning Sparse Instance-Dependent Attention for Efficient Vision TransformersabstractVision Transformers (ViT) have shown competitive advantages in terms of performance compared to convolutional neural networks (CNNs), though they often come with high computational costs. To this end, previous methods explore different attention patterns by limiting a fixed number of spatially nearby tokens to accelerate the ViT's multi-head self-attention (MHSA) operations. However, such structured attention patterns limit the token-to-token connections to their spatial relevance, which disregards learned semantic connections from a full attention mask. In this work, we propose an approach to learn instance-dependent attention patterns, by devising a lightweight connectivity predictor module that estimates the connectivity score of each pair of tokens. Intuitively, two tokens have high connectivity scores if the features are considered relevant either spatially or semantically. As each token only attends to a small number of other tokens, the binarized connectivity masks are often very sparse by nature and therefore provide the opportunity to reduce network FLOPs via sparse computations. Equipped with the learned unstructured attention pattern, sparse attention ViT (Sparsifiner) produces a superior Pareto frontier between FLOPs and top-1 accuracy on ImageNet compared to token sparsity. Our method reduces 48% ~ 69% FLOPs of MHSA while the accuracy drop is within 0.4%. We also show that combining attention and token sparsity reduces ViT FLOPs by over 60%. Cong Wei 0001, Brendan Duke, Ruowei Jiang, Parham Aarabi, Graham W. Taylor, Florian Shkurti |
CVPR | 3 |
| 2022 | Exploring Gradient-Based Multi-directional Controls in GANs
Ruowei Jiang, Brendan Duke, Han Zhao 0002, Parham Aarabi |
ECCV (23) | 2 |
| 2022 | Synthesizing ultraviolet skin images via GAN with Gaussian weighted patch blendingabstractIn this work, we explore a novel application of synthesizing ultraviolet skin images from RGB images using an unpaired training framework for image-to-image translation. To synthesize high resolution outputs, we propose a novel Gaussian-based patch blending technique that is designed following the characteristics of GANs. Specifically, we weigh the pixels at the same coordinates among multiple generated patches based on their distance to the center point. Our proposed method is performant, taking 0.72s for a whole face image at resolution 960x720 at inference time and generates realistic-looking ultraviolet images. We also show high correspondence of our synthesized images with the true ultraviolet images qualitatively. Finally, our novel blending approach achieves significant improvements compared with other blending methods. Ruowei Jiang, Brendan Duke, Frédéric Flament, Parham Aarabi |
ISM | 1 |
| 2022 | Real-time Virtual-Try-On from a Single Example Image through Deep Inverse Graphics and Learned Differentiable RenderersabstractAbstract Augmented reality applications have rapidly spread across online retail platforms and social media, allowing consumers to virtually try‐on a large variety of products, such as makeup, hair dying, or shoes. However, parametrizing a renderer to synthesize realistic images of a given product remains a challenging task that requires expert knowledge. While recent work has introduced neural rendering methods for virtual try‐on from example images, current approaches are based on large generative models that cannot be used in real‐time on mobile devices. This calls for a hybrid method that combines the advantages of computer graphics and neural rendering approaches. In this paper, we propose a novel framework based on deep learning to build a real‐time inverse graphics encoder that learns to map a single example image into the parameter space of a given augmented reality rendering engine. Our method leverages self‐supervised learning and does not require labeled training data, which makes it extendable to many virtual try‐on applications. Furthermore, most augmented reality renderers are not differentiable in practice due to algorithmic choices or implementation constraints to reach real‐time on portable devices. To relax the need for a graphics‐based differentiable renderer in inverse graphics problems, we introduce a trainable imitator module. Our imitator is a generative network that learns to accurately reproduce the behavior of a given non‐differentiable renderer. We propose a novel rendering sensitivity loss to train the imitator, which ensures that the network learns an accurate and continuous representation for each rendering parameter. Automatically learning a differentiable renderer, as proposed here, could be beneficial for various inverse graphics tasks. Our framework enables novel applications where consumers can virtually try‐on a novel unknown product from an inspirational reference image on social media. It can also be used by computer graphics artists to automatically create realistic rendering from a reference product image. Robin Kips, Ruowei Jiang, Sileye O. Ba, Brendan Duke, Matthieu Perrot, Pietro Gori, Isabelle Bloch |
Comput. Graph. Forum | 2 |
| 2021 | Continuous Face Aging via Self-Estimated Residual Age EmbeddingabstractFace synthesis, including face aging, in particular, has been one of the major topics that witnessed a substantial improvement in image fidelity by using generative adversarial networks (GANs). Most existing face aging approaches divide the dataset into several age groups and leverage group-based training strategies, which lacks the ability to provide fine-controlled continuous aging synthesis in nature. In this work, we propose a unified network structure that embeds a linear age estimator into a GAN-based model, where the embedded age estimator is trained jointly with the encoder and decoder to estimate the age of a face image and provide a personalized target age embedding for age progression/regression. The personalized target age embedding is synthesized by incorporating both personalized residual age embedding of the current age and exemplar-face aging basis of the target age, where all preceding aging bases are derived from the learned weights of the linear age estimator. This formulation brings the unified perspective of estimating the age and generating personalized aged face, where self-estimated age embeddings can be learned for every single age. The qualitative and quantitative evaluations on different datasets further demonstrate the significant improvement in the continuous face aging aspect over the state-of-the-art. Zeqi Li, Ruowei Jiang, Parham Aarabi |
CVPR | 2 |
| 2020 | Semantic Relation Preserving Knowledge Distillation for Image-to-Image Translation
Zeqi Li, Ruowei Jiang, Parham Aarabi |
ECCV (26) | 2 |