EDBT 2026 Demo / reviewers in the wild / expert
Faizan Farooq Khan
dblp:302/1121
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0002-3645-603XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Step-by-step Layered Design Generation
Faizan Farooq Khan, K. J. Joseph, Koustava Goswami, Mohamed Elhoseiny 0001, Balaji Vasan Srinivasan |
AAAI | 1 |
| 2026 | Sketch2Stitch: GANs for Abstract Sketch-Based Dress SynthesisabstractIn the realm of creative expression, not everyone possesses the gift of effortlessly translating their imaginative visions into flawless sketches. More often than not, the outcome resembles an abstract, perhaps even slightly distorted representation. The art of producing impeccable sketches is not only challenging but also a time-consuming process. Our work is the first of this kind in transforming abstract, sometimes deformed garment sketches into photorealistic catalog images, to empower the everyday individual to become their own fashion designer. We create Sketch2Stitch, a dataset featuring over 65,000 abstract sketch images generated from garments of Dress Code [40] and VITON HD [6], two benchmark datasets in the virtual try-on task. Sketch2Stitch is the first dataset in the literature to provide abstract sketches in the fashion domain. We propose a StyleGAN-based generative framework that bridges freehand sketching with photorealistic garment synthesis. We demonstrate that our framework allows users to sketch rough outlines and optionally provide color hints, producing realistic designs in seconds. Experimental results demonstrate, both quantitatively and qualitatively, that the proposed framework achieves superior performance against various baselines and existing methods on both subsets of our dataset. Our work highlights a pathway toward AI-assisted fashion design tools, democratizing garment ideation for students, independent designers, and casual creators. Faizan Farooq Khan, Eslam Abdelrahman, Davide Morelli, Marcella Cornia, Rita Cucchiara, Mohamed Elhoseiny 0001 |
WACV | 1 |
| 2025 | Local Masked Reconstruction for Efficient Self-Supervised Learning on High-Resolution ImagesabstractSelf-supervised learning for computer vision has progressed tremendously and improved many downstream vision tasks, such as image classification, semantic segmentation, and object detection. Among these, generative self-supervised vision learning approaches, such as MAE and BEiT, show promising performance. However, their global reconstruction mechanism is computationally demanding, especially for high-resolution images. The computational cost increases extensively when scaled to a large-scale dataset. To address this issue, we propose local masked reconstruction (LoMaR), a simple yet effective approach that reconstructs image patches from small neighboring regions. The strategy can be easily integrated into any generative self-supervised learning techniques and improves the trade-off between efficiency and accuracy compared to reconstruction over the entire image. LoMaR is$2.5\times faster$than MAE and 5.0x faster than BEiT on$384\times 384$ImageNet pretraining and surpasses them by 0.2% and 0.8% in accuracy, respectively. It is$2.1\times faster$than MAE on iNaturalist pretraining and gains 0.2% in accuracy. On MS COCO, LoMaR outperforms MAE by 0.5$AP^{box}$on object detection and 0.5$AP^{mask}$on instance segmentation. It also outperforms$MAE$by 0.2% on semantic segmentation. Our code and pretrained models are available at: https://github.com/junchen14/LoMaR. Jun Chen 0021, Faizan Farooq Khan, Ammar Sherif, ZongYuan Ge, Boyang Li 0001, Mohamed Elhoseiny 0001 |
WACV | 2 |
| 2024 | EmoTalker: Audio Driven Emotion Aware Talking Head Generation
Xiaoqian Shen, Faizan Farooq Khan, Mohamed Elhoseiny 0001 |
ACCV (5) | 2 |
| 2023 | HRS-Bench: Holistic, Reliable and Scalable Benchmark for Text-to-Image ModelsabstractIn recent years, Text-to-Image (T2I) models have been extensively studied, especially with the emergence of diffusion models that achieve state-of-the-art results on T2I synthesis tasks. However, existing benchmarks heavily rely on subjective human evaluation, limiting their ability to holistically assess the model’s capabilities. Furthermore, there is a significant gap between efforts in developing new T2I architectures and those in evaluation. To address this, we introduce HRS-Bench, a concrete evaluation benchmark for T2I models that is Holistic, Reliable, and Scalable. Unlike existing benchmarks that focus on limited aspects, HRS-Bench measures 13 skills that can be categorized into five major categories: accuracy, robustness, generalization, fairness, and bias. In addition, HRS-Bench covers 50 scenarios, including fashion, animals, transportation, food, and clothes. We evaluate nine recent large-scale T2I models using metrics that cover a wide range of skills. A human evaluation aligned with 95% of our evaluations on average was conducted to probe the effectiveness of HRS-Bench. Our experiments demonstrate that existing models often struggle to generate images with the desired count of objects, visual text, or grounded emotions. We hope that our benchmark help ease future text-to-image generation research. The code and data are available at https://eslambakr.github.io/hrsbench.github.io/. Eslam Mohamed Bakr, Pengzhan Sun 0001, Xiaoqian Shen, Faizan Farooq Khan, Li Erran Li, Mohamed Elhoseiny 0001 |
ICCV | 4 |
| 2023 | FishNet: A Large-scale Dataset and Benchmark for Fish Recognition, Detection, and Functional Trait PredictionabstractAquatic species are essential components of the world’s ecosystem, and the preservation of aquatic biodiversity is crucial for maintaining proper ecosystem functioning. Unfortunately, increasing anthropogenic pressures such as overfishing, climate change, and coastal development pose significant threats to aquatic biodiversity. To address this challenge, it is necessary to design an automatic aquatic species monitoring systems that can help researchers and policymakers better understand changes in aquatic ecosystems and take appropriate actions to preserve biodiversity. However, the development of such systems is impeded by a lack of large-scale diverse aquatic species datasets. Existing aquatic species recognition datasets generally have a limited number of species, nor do they provide functional trait data, and so have only narrow potential for application. To address the need for generalized systems that can recognize, locate, and predict a wide array of species and their functional traits, we present FishNet, a large-scale diverse dataset containing 94,532 meticulously organized images from 17,357 aquatic species, organized according to aquatic biological taxonomy (order, family, genus, and species). We further build three benchmarks, i.e., fish classification, fish detection, and functional trait prediction, inspired by ecological research needs, to facilitate the development of aquatic species recognition systems, and promote further research in the field of aquatic ecology. Our FishNet dataset has the potential to encourage the development of more accurate and effective tools for the monitoring and protection of aquatic ecosystems, and hence take effective action toward the conservation of our planet’s aquatic biodiversity. Our dataset and code will be released at https://fishnet-2023.github.io/. Faizan Farooq Khan, Xiang Li 0046, Andrew J. Temple, Mohamed Elhoseiny 0001 |
ICCV | 1 |
| 2023 | Towards Generating Ultra-High Resolution Talking-Face Videos with Lip synchronizationabstractTalking-face video generation works have achieved state-of-the-art results in synthesizing videos with lip synchronization. However, most of the previous works deal with low-resolution talking-face videos (up to 256×256 pixels), thus, generating extremely high-resolution videos still remains a challenge. We take a giant leap in this work and propose a novel method to synthesize talking-face videos at resolutions as high as 4K! Our task presents several key challenges: (i) Scaling the existing methods to such high resolutions is resource-constrained, both in terms of compute and the availability of very high-resolution datasets, (ii) The synthesized videos need to be spatially and temporally coherent. The sheer number of pixels that the model needs to generate while maintaining the temporal consistency at the video level makes this task non-trivial and has never been attempted before in literature. To address these issues, we propose to train the lip-sync generator in a compact Vector Quantized (VQ) space for the first time. Our core idea to encode the faces in a compact 16× 16 representation allows us to model high-resolution videos. In our framework, we learn the lip movements in the quantized space on the newly collected 4K Talking Faces (4KTF) dataset. Our approach is speaker agnostic and can handle various languages and voices. We benchmark our technique against several competitive works and show that we can achieve a remarkable 64-times more pixels than the current state-of-the-art! Our supplementary demo video depicts additional qualitative results, comparisons, and several real-world applications, like professional movie editing enabled by our model. Anchit Gupta, Rudrabha Mukhopadhyay, Sindhu Balachandra, Faizan Farooq Khan, Vinay P. Namboodiri, C. V. Jawahar |
WACV | 4 |
| 2022 | It is Okay to Not Be Okay: Overcoming Emotional Bias in Affective Image Captioning by Contrastive Data CollectionabstractDatasets that capture the connection between vision, language, and affection are limited, causing a lack of understanding of the emotional aspect of human intelligence. As a step in this direction, the ArtEmis dataset was recently introduced as a large-scale dataset of emotional reactions to images along with language explanations of these chosen emotions. We observed a significant emotional bias towards instance-rich emotions, making trained neural speakers less accurate in describing under-represented emotions. We show that collecting new data, in the same way, is not effective in mitigating this emotional bias. To remedy this problem, we propose a contrastive data collection approach to balance ArtEmis with a new complementary dataset such that a pair of similar images have contrasting emotions (one positive and one negative). We collected 260,533 instances using the proposed method, we combine them with ArtEmis, creating a second iteration of the dataset. The new combined dataset, dubbed ArtEmis v2.0, has a balanced distribution of emotions with explanations revealing more fine details in the associated painting. Our experiments show that neural speakers trained on the new dataset improve CIDEr and METEOR evaluation metrics by 20% and 7%, respectively, compared to the biased dataset. Finally, we also show that the performance per emotion of neural speakers is improved across all the emotion categories, significantly on under-represented emotions. The collected dataset and code are available at https://artemisdataset-v2.org. Youssef Mohamed, Faizan Farooq Khan, Kilichbek Haydarov, Mohamed Elhoseiny 0001 |
CVPR | 2 |