Vinay Kaushik

dblp:232/1964 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0001-6729-4857ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 DSFace : Conditional Diffusion Inpainting for Sketch-to-Face Synthesis
abstract
Generating realistic human faces from monochromatic sketches is a challenging task with applications in forensic reconstruction, character design, and digital art. The limited semantic information in single-channel sketches makes this problem difficult, as they often lack fine details like expressions, skin tone and accessories. While GANs have shown promise, they suffer from unstable training and poor structural guidance. Diffusion models offer improved image generation, but struggle with monochrome inputs and high computational costs. To address these challenges, we propose DSFace, a latent diffusion-based framework that treats sketch-to-face generation as a conditional inpainting problem. DSFace utilises a frozen Paint-by-Example (PBE) inpainting diffusion model, conditioned with a ControlNet encoder, ensuring precise control over face synthesis. Our novel approach utilises a GAN-generated coarse image to compute DINO-V2 embeddings, which provide fine-grained details for improved facial and garment features. Trained on the CUFS dataset, DSFace achieves state-of-the-art performance, surpassing existing methods in visual realism, perceptual quality and structural alignment with the input sketches.
Sanhita Pathak, Vinay Kaushik, Brejesh Lall
ICIP2
2025 CLEARSTR: Contextual Learning with Edge-guided and Adaptive-texture Reconstruction for Scene Text Removal
abstract
Scene text removal is a challenging task in computer vision, requiring the seamless restoration of text-masked regions while preserving the structural and aesthetic coherence of the background. Current methods often fail to achieve natural integration of the restored regions into the surrounding context, especially in complex scenes. We introduce structure guidance to the task of scene text removal utilising a novel framework that extends denoising diffusion probabilistic models (DDPMs) for text removal tasks. Our approach integrates depth-aware neighborhood estimation to identify regions with similar depth profiles near the text-masked area, providing spatial cues to guide the inpainting process. Additionally, the model leverages localized texture reconstruction, ensuring that the synthesized textures align with the intricate details of the surrounding image. We propose a unified approach for context-aware guidance that dynamically integrates both depth and spatial proximity constraints into a single, coherent neighborhood definition. To ensure semantic consistency in generated scene image we also propose context loss. We evaluate our approach on the SCUT-EnsText and SCUT-Syn datasets, demonstrating its ability to achieve superior text removal quality, combining high perceptual fidelity with robust quantitative performance. By incorporating structural depth information and context-aware texture generation, this work sets a new benchmark in scene text removal research.
Sanhita Pathak, Vinay Kaushik, Brejesh Lall
ICME2
2025 COVITON : Consistency driven integration of TPS and flow for virtual tryon
Sanhita Pathak, Vinay Kaushik, Brejesh Lall
Comput. Graph.2
2025 Garment Recycle Training and Conditional Garment-Person Outline Attention-Guided Virtual Tryon
abstract
Virtual try-on, a significant application in computer vision, aims to seamlessly simulate the appearance of clothing on a person from a single image. We propose a diffusion-based tryon approach, solving virtual tryon as a problem of conditional image inpainting. Our method introduces GarNet and OutlineNet as two learnable Stable Diffusion ControlNet encoders conditioned on the garment and person outline images, enhancing the controllability and realism of the generated try-on. We propose a two-stage garment diffusion recycling training strategy, utilizing \(x_{0}\) -parameterization. We estimate the initial clean image that is conditioned on the maximum noised input and feed the same to the same diffusion model again to estimate total noise. This reduces over-fitting and makes our model more generalized. We also introduce a zero garment-outline conditioning (ZGOC) block along with a Garment-Outline Cross Attention layer to optimize garment draping and ensure global consistency in the try-on results. The ZGOC block provides control and adaptability by prioritizing garment details that are most affected by body shape, ensuring precise garment alignment with the person’s outline. Our comprehensive experiments on the VITON-HD and Dresscode dataset demonstrate that our proposed approach achieves state-of-the-art realism and controllability in VITON, marking a significant advancement in virtual fashion experiences and online shopping applications.
Sanhita Pathak, Vinay Kaushik, Brejesh Lall
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Single Stage Warped Cloth Learning and Semantic-Contextual Attention Feature Fusion for Virtual Tryon
abstract
Image-based virtual try-on aims to fit an in-shop garment onto a clothed person image. Garment warping, which aligns the target garment with the corresponding body parts in the person image, is a crucial step in achieving this goal. Existing methods often use multi-stage frameworks to handle clothes warping, person body synthesis and tryon generation separately or rely on noisy intermediate parser-based labels. We propose a novel single-stage framework that implicitly learns the same without explicit multi-stage learning. Our approach utilizes a novel semantic-contextual fusion attention module for garment-person feature fusion, enabling efficient and realistic cloth warping and body synthesis from target pose keypoints. By introducing a lightweight linear attention framework that attends to garment regions and fuses multiple sampled flow fields, we also address misalignment and artifacts present in previous methods. To achieve simultaneous learning of warped garment and try-on results, we introduce a Warped Cloth Learning Module. Our proposed approach significantly improves the quality and efficiency of virtual try-on methods, providing users with a more reliable and realistic virtual try-on experience.
Sanhita Pathak, Vinay Kaushik, Brejesh Lall
ICME2
2024 ICPR 2024 Leaf Inspect Competition: Leaf Instance Segmentation and Counting
Swati Bhugra, Prerana Mukherjee, Vinay Kaushik, Siddharth Srivastava 0004, Viswanathan Chinnusamy, Brejesh Lall, Santanu Chaudhary
ICPR (34)3
2023 Hierarchical Multi-task Learning via Task Affinity Groupings
abstract
Multi-task learning (MTL) permits joint task learning based on a shared deep learning architecture and multiple loss functions. Despite the recent advances in MTL, one loss often dominates the learning optimization in multiple unrelated tasks. This often results in poor performance compared to the corresponding single task learning. To overcome the aforementioned "negative transfer", we propose a novel hierarchical framework that leverages task relations via inter-task affinity to supervise multi-task learning. Specifically, the inter-task affinity generated task sets, with low-level task set and complex task set at the bottom and top layers respectively, enables iterative multi-task information sharing. In addition, it also alleviates simultaneous image annotations for multiple tasks. The proposed framework achieves state-of-the-art results on classification, detection, semantic segmentation and depth estimation across three standard benchmarks. Furthermore, with state of the results on two benchmarks for image retrieval task, we also demonstrate that the embeddings learned using such a framework provide good generalization and robust representation learning.
Siddharth Srivastava 0004, Swati Bhugra, Vinay Kaushik, Brejesh Lall
ICIP3
2023 AnoLeaf: Unsupervised Leaf Disease Segmentation via Structurally Robust Generative Inpainting
abstract
Plant diseases severely limits agriculture production, necessitating the high-throughput monitoring of plant leaves. Currently, this is formulated as an automatic disease segmentation task addressed via deep learning frameworks. These deep leaning frameworks trained with leaf image data in a supervised paradigm have few limitations, mainly: (1) training datasets are heavily imbalanced towards healthy leaf images, (2) disease region annotation is labour-intensive and (3) due to the heterogeneity of disease symptoms, these frameworks lacks generalisability. In this paper, we reformulate disease segmentation as an anomaly localisation task. Specifically, we introduce a novel unsupervised framework (AnoLeaf) based on an edge-guided in-painting that optimises the learning of contextual attention on only healthy leaf images. The network utilisation on diseased leaf images results in reconstruction of its healthy counterparts, generating an inpainting error. The contextual attention maps reinforce the inpainting error to effectively localise the disease. Thus, AnoLeaf alleviates the acquisition and annotation of rare disease images. Additional experiments on MVTec anomaly detection dataset further demonstrate its generalisability.
Swati Bhugra, Vinay Kaushik, Brejesh Lall, Santanu Chaudhury
WACV2
2019 UnDispNet: Unsupervised Learning for Multi-Stage Monocular Depth Prediction
abstract
Despite advances in single view depth estimation, most existing techniques treat the task in a supervised manner. Recent approaches utilize the possibility of learning without ground truth depth, by minimizing the photometric error. In this paper, we propose a deep framework that refines predicted depth from a single image using a two-stage process, exploiting sub-pixel convolutions for depth super resolution. The first stage uses a pyramidal input and learns depth at 4 scales, utilizing depth super resolution. The second stage uses warp errors, reconstructed images, predicted depth along with original left input to refine the depth predicted by the first stage. We use data augmentation by varying color, scale and incorporating left-right flipped images in our data. We train our network in a completely unsupervised way on photoconsistency imposing occlusion, left right consistency and disparity smoothness constraints. We transform the learning process into optimally distributed steps, varying the combination of scales and losses to minimize over-regularization of depth maps. We evaluate our model for monocular inputs on KITTI driving benchmark. Our depth predictions surpass state-of-the art self-supervised approaches for monocular depth prediction.
Vinay Kaushik, Brejesh Lall
3DV1