Maitreya Suin

dblp:254/1457 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0002-0004-181XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 9 · 5 first-author · 7 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Improved Representation Learning for Unconstrained Face Recognition
abstract
Face recognition is a widely studied problem where the aim is to design a robust network that assigns higher similarity to the same face and reduces similarity between dissimilar faces. Previous research utilizing margin-based loss functions has achieved near-perfect accuracies on high-quality face recognition datasets. However, the same networks fail to perform well on low-quality images due to the degradation of facial attributes necessary for distinguishing different faces. In this paper, we tackle the problem of low-quality face recognition. We base our analysis on an observation that the change of loss functions produce marginal changes in performance for low-quality face recognition. Hence, rather than following the traditional approach of defining problem-specific regularized functions, we take a closer look at the nature of data in low resolution datasets and redefine paradigms in terms of model choice, data input pipeline and fine-tuning schemes. With the accumulated effect of all our design choices, we achieve state-of-the-art results in medium-quality benchmarks (IJB-B, IJB-C) as well as multiple challenging benchmarks for unconstrained face recognition (Tinyface, IJB-S and BRIAR), thereby opening up a new avenue of research in the area. The pretrained model are publically available in https://github.com/ Kartik-3004/PETALface
Nithin Gopalakrishnan Nair, Kartik Narayan, Maitreya Suin, Ram Prabhakar Kathirvel, Soraya Stevens, Joshua Gleason, Nathan Shnidman, Rama Chellappa, Vishal M. Patel
FG3
2024 CLR-Face: Conditional Latent Refinement for Blind Face Restoration Using Score-Based Diffusion Models
Maitreya Suin, Rama Chellappa
IJCAI1
2024 Diffuse and Restore: A Region-Adaptive Diffusion Model for Identity-Preserving Blind Face Restoration
abstract
Blind face restoration (BFR) from severely degraded face images in the wild is a highly ill-posed problem. Due to the complex unknown degradation, existing generative works typically struggle to restore realistic details when the input is of poor quality. Recently, diffusion-based approaches were successfully used for high-quality image synthesis. But, for BFR, maintaining a balance between the fidelity of the restored image and the reconstructed identity information is important. Minor changes in certain facial regions may alter the identity or degrade the perceptual quality. With this observation, we present a conditional diffusion-based framework for BFR. We alleviate the drawbacks of existing diffusion-based approaches and design a region-adaptive strategy. Specifically, we use an identity preserving conditioner network to recover the identity information from the input image as much as possible and use that to guide the reverse diffusion process, specifically for important facial locations that contribute the most to the identity. This leads to a significant improvement in perceptual quality as well as face-recognition scores over existing GAN and diffusion-based restoration models. Our approach achieves superior results to prior art on a range of real and synthetic datasets, particularly for severely degraded face images.
Maitreya Suin, Nithin Gopalakrishnan Nair, Chun Pong Lau 0001, Vishal M. Patel, Rama Chellappa
WACV1
2023 Exploring the Effectiveness of Mask-Guided Feature Modulation as a Mechanism for Localized Style Editing of Real Images (Student Abstract)
abstract
The success of Deep Generative Models at high-resolution image generation has led to their extensive utilization for style editing of real images. Most existing methods work on the principle of inverting real images onto their latent space, followed by determining controllable directions. Both inversion of real images and determination of controllable latent directions are computationally expensive operations. Moreover, the determination of controllable latent directions requires additional human supervision. This work aims to explore the efficacy of mask-guided feature modulation in the latent space of a Deep Generative Model as a solution to these bottlenecks. To this end, we present the SemanticStyle Autoencoder (SSAE), a deep Generative Autoencoder model that leverages semantic mask-guided latent space manipulation for highly localized photorealistic style editing of real images. We present qualitative and quantitative results for the same and their analysis. This work shall serve as a guiding primer for future work.
Snehal Singh Tomar, Maitreya Suin, A. N. Rajagopalan 0001
AAAI2
2023 ATDetect: Face Detection and Keypoint Extraction at Range and Altitude
abstract
Face detection and alignment are the crucial preprocessing steps in face recognition. While face detection works well in ideal situations, the performance deteriorates significantly when the image is degraded, due to factors such as blur, deformation, low resolution, and extreme headpose. However, there are very few works on face detection with realistic data captured from a long range (100m to 500m) and high altitude (30° to 50° pitch angle). We first evaluated several state-of-the-art methods on data collected at ranges of 100-500 meters and large pitch angles. One challenge is videos captured from long ranges usually lack bounding boxes and keypoint annotations, needed for training deep networks. This motivates us to develop a face detection and alignment algorithm that could perform effectively on videos captured from a long range and high altitude without groundtruth annotations. Moreover, meta information such as age, gender, and headpose of the subject could help face recognition. Therefore, we propose a single-stage face localization model ATDetect, which detects face bounding boxes, keypoints, and meta information simultaneously with realistic video captured at range and altitude.
Chun Pong Lau 0001, Maitreya Suin, Rama Chellappa
IJCB2
2023 Illumination-Adaptive Unpaired Low-Light Enhancement
abstract
Supervised networks address the task of low-light enhancement using paired images. However, collecting a wide variety of low-light/clean paired images is tedious as the scene needs to remain static during imaging. In this paper, we propose an unsupervised low-light enhancement network using context-guided illumination-adaptive norm (CIN). Inspired by coarse to fine methods, we propose to address this task in two stages. In stage- I, a pixel amplifier module (PAM) is used to generate a coarse estimate with an overall improvement in visibility and aesthetic quality. Stage- II further enhances the saturated dark pixels and scene properties of the image using CIN. Different ablation studies show the importance of PAM and CIN in improving the visible quality of the image. Next, we propose a region-adaptive single input multiple output (SIMO) model that can generate multiple enhanced images from a single low-light image. The objective of SIMO is to let users choose the image of their liking from a pool of enhanced images. Human subjective analysis of SIMO results shows that the distribution of preferred images varies, endorsing the importance of SIMO-type models. Lastly, we propose a low-light road scene (LLRS) dataset having an unpaired collection of low-light and clean scenes. Unlike existing datasets, the clean and low-light scenes in LLRS are real and captured using fixed camera settings. Exhaustive comparisons on publicly available datasets, and the proposed dataset reveal that the results of our model outperform prior art quantitatively and qualitatively.
Praveen Kandula, Maitreya Suin, A. N. Rajagopalan 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 Gated Spatio-Temporal Attention-Guided Video Deblurring
abstract
Video deblurring remains a challenging task due to the complexity of spatially and temporally varying blur. Most of the existing works depend on implicit or explicit alignment for temporal information fusion, which either increases the computational cost or results in suboptimal performance due to misalignment. In this work, we investigate two key factors responsible for deblurring quality: how to fuse spatio-temporal information and from where to collect it. We propose a factorized gated spatio-temporal attention module to perform non-local operations across space and time to fully utilize the available information without depending on alignment. First, we perform spatial aggregation followed by a temporal aggregation step. Next, we adaptively distribute the global spatio-temporal information to each pixel. It shows superior performance compared to existing non-local fusion techniques while being considerably more efficient. To complement the attention module, we propose a reinforcement learning-based framework for selecting keyframes from the neighborhood with the most complementary and useful information. Moreover, our adaptive approach can increase or decrease the frame usage at inference time, depending on the user’s need. Extensive experiments on multiple datasets demonstrate the superiority of our method.
Maitreya Suin, A. N. Rajagopalan 0001
CVPR1
2021 Spatially-Adaptive Image Restoration using Distortion-Guided Networks
abstract
We present a general learning-based solution for restoring images suffering from spatially-varying degradations. Prior approaches are typically degradation-specific and employ the same processing across different images and different pixels within. However, we hypothesize that such spatially rigid processing is suboptimal for simultaneously restoring the degraded pixels as well as reconstructing the clean regions of the image. To overcome this limitation, we propose SPAIR, a network design that harnesses distortion-localization information and dynamically adjusts computation to difficult regions in the image. SPAIR comprises of two components, (1) a localization network that identifies degraded pixels, and (2) a restoration network that exploits knowledge from the localization network in filter and feature domain to selectively and adaptively restore degraded pixels. Our key idea is to exploit the non-uniformity of heavy degradations in spatial-domain and suitably embed this knowledge within distortion-guided modules performing sparse normalization, feature extraction and attention. Our architecture is agnostic to physical formation model and generalizes across several types of spatially-varying degradations. We demonstrate the efficacy of SPAIR individually on four restoration tasks- removal of rain-streaks, raindrops, shadows and motion blur. Extensive qualitative and quantitative comparisons with prior art on 11 benchmark datasets demonstrate that our degradation-agnostic network design offers significant performance gains over state-of-the-art degradation-specific architectures. Code available at https://github.com/humananalysis/spatially-adaptive-image-restoration.
Kuldeep Purohit, Maitreya Suin, A. N. Rajagopalan 0001, Vishnu Naresh Boddeti
ICCV2
2021 Distillation-guided Image Inpainting
abstract
Image inpainting methods have shown significant improvements by using deep neural networks recently. However, many of these techniques often create distorted structures or blurry inconsistent textures. The problem is rooted in the encoder layers’ ineffectiveness in building a complete and faithful embedding of the missing regions from scratch. Existing solutions like course-to-fine, progressive refinement, structural guidance, etc. suffer from huge computational overheads owing to multiple generator networks, limited ability of handcrafted features, and sub-optimal utilization of the information present in the ground truth. We propose a distillation-based approach for inpainting, where we provide direct feature level supervision while training. We deploy cross and self-distillation techniques and design a dedicated completion-block in encoder to produce more accurate encoding of the holes. Next, we demonstrate how an inpainting network’s attention module can improve by leveraging a distillation-based attention transfer technique and further enhance coherence by using a pixeladaptive global-local feature fusion. We conduct extensive evaluations on multiple datasets to validate our method. Along with achieving significant improvements over previous SOTA methods, the proposed approach’s effectiveness is also demonstrated through its ability to improve existing inpainting works.
Maitreya Suin, Kuldeep Purohit, A. N. Rajagopalan 0001
ICCV1
2020 An Efficient Framework for Dense Video Captioning
abstract
Dense video captioning is an extremely challenging task since an accurate and faithful description of events in a video requires a holistic knowledge of the video contents as well as contextual reasoning of individual events. Most existing approaches handle this problem by first proposing event boundaries from a video and then captioning on a subset of the proposals. Generation of dense temporal annotations and corresponding captions from long videos can be dramatically source consuming. In this paper, we focus on the task of generating a dense description of temporally untrimmed videos and aim to significantly reduce the computational cost by processing fewer frames while maintaining accuracy. Existing video captioning methods sample frames with a predefined frequency over the entire video or use all the frames. Instead, we propose a deep reinforcement-based approach which enables an agent to describe multiple events in a video by watching a portion of the frames. The agent needs to watch more frames when it is processing an informative part of the video, and skip frames when there is redundancy. The agent is trained using actor-critic algorithm, where the actor determines the frames to be watched from a video and the critic assesses the optimality of the decisions taken by the actor. Such an efficient frame selection simplifies the event proposal task considerably. This has the added effect of reducing the occurrence of unwanted proposals. The encoded state representation of the frame selection agent is further utilized for guiding event proposal and caption generation tasks. We also leverage the idea of knowledge distillation to improve the accuracy. We conduct extensive evaluations on ActivityNet captions dataset to validate our method.
Maitreya Suin, A. N. Rajagopalan 0001
AAAI1
2020 Spatially-Attentive Patch-Hierarchical Network for Adaptive Motion Deblurring
abstract
This paper tackles the problem of motion deblurring of dynamic scenes. Although end-to-end fully convolutional designs have recently advanced the state-of-the-art in non-uniform motion deblurring, their performance-complexity trade-off is still sub-optimal. Existing approaches achieve a large receptive field by increasing the number of generic convolution layers and kernel-size, but this comesat the expense of of the increase in model size and inference speed. In this work, we propose an efficient pixel adaptive and feature attentive design for handling large blur variations across different spatial locations and process each test image adaptively. We also propose an effective content-aware global-local filtering module that significantly improves performance by considering not only global dependencies but also by dynamically exploiting neighboring pixel information. We use a patch-hierarchical attentive architecture composed of the above module that implicitly discovers the spatial variations in the blur present in the input image and in turn, performs local and global modulation of intermediate features. Extensive qualitative and quantitative comparisons with prior art on deblurring benchmarks demonstrate that our design offers significant improvements over the state-of-the-art in accuracy as well as speed.
Maitreya Suin, Kuldeep Purohit, A. N. Rajagopalan 0001
CVPR1