Shanmuganathan Raman

dblp:70/4688 · DBLP profile ↗
← Back
58ranked-venue papers
3as first author
35since 2021 · last 2026
0000-0003-2718-7891ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 50 · 3 first-author · 33 since 2021Artificial intelligence and machine learning · 20 · 1 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
YearPublicationVenuePosition
2026 Synthesizing Compositional Videos from Text Description
abstract
Existing pre-trained text-to-video diffusion models can generate high-quality videos, but often struggle with misalignment between the generated content and the input text, particularly while composing scenes with multiple objects. To tackle this issue, we propose a straightforward, training-free approach for compositional video generation from text. We introduce Video-ASTAR for test-time aggregation and segregation of attention with a novel centroid loss to enhance alignment, which enables the generation of multiple objects in the scene, modeling the actions and interactions. Additionally, we extend our approach to the Multi-Action video generation setting, where only the specified action should vary across a sequence of prompts. To ensure coherent action transitions, we introduce a novel token-swapping and latent interpolation strategy. Extensive experiments and ablation studies show that our method significantly outperforms baseline methods, generating videos with improved semantic and compositional consistency alongside improved temporal coherence1.
Prajwal Singh, Kuldeep Kulkarni, Shanmuganathan Raman, Harsh Rangwani
WACV3
2026 GenR: Generative latent inversion for blind face restoration
Indra Deep Mastan, Shanmuganathan Raman
Pattern Recognit. Lett.3
2025 RASP: Revisiting 3D Anamorphic Art for Shadow-Guided Packing of Irregular Objects
abstract
Recent advancements in learning-based methods have opened new avenues for exploring and interpreting art forms, such as shadow art, origami, and sketch art, through computational models. One notable visual art form is 3D Anamorphic Art in which an ensemble of arbitrarily shaped 3D objects creates a realistic and meaningful expression when observed from a particular viewpoint and loses its coherence over the other viewpoints. In this work, we build on insights from 3D Anamorphic Art to perform 3D object arrangement. We introduce RASP, a differentiable-rendering-based framework to arrange arbitrarily shaped 3D objects within a bounded volume via shadow (or silhouette)-guided optimization with an aim of minimal inter-object spacing and near-maximal occupancy. Furthermore, we propose a novel SDF-based formulation to handle inter-object intersection and container extrusion. We demonstrate that RASP can be extended to part assembly alongside object packing considering 3D objects to be "parts" of another 3D object. Finally, we present artistic illustrations of multi-view anamorphic art, achieving meaningful expressions from multiple viewpoints within a single ensemble.
Soumyaratna Debnath, Ashish Tiwari 0005, Kaustubh Sadekar, Shanmuganathan Raman
CVPR4
2025 L3D-Pose: Lifting Pose for 3D Avatars from a Single Camera in the Wild
abstract
While 2D pose estimation has advanced our ability to interpret body movements in animals and primates, it is limited by the lack of depth information, constraining its application range. 3D pose estimation provides a more comprehensive solution by incorporating spatial depth, yet creating extensive 3D pose datasets for animals is challenging due to their dynamic and unpredictable behaviours in natural settings. To address this, we propose a hybrid approach that utilizes rigged avatars and the pipeline to generate synthetic datasets to acquire the necessary 3D annotations for training. Our method introduces a simple attention-based MLP network for converting 2D poses to 3D, designed to be independent of the input image to ensure scalability for poses in natural environments. Additionally, we identify that existing anatomical keypoint detectors are insufficient for accurate pose retargeting onto arbitrary avatars. To overcome this, we present a lookup table based on a deep pose estimation method using a synthetic collection of diverse actions rigged avatars perform. Our experiments demonstrate the effectiveness and efficiency of this lookup table-based retargeting approach. Overall, we propose a comprehensive framework with systematically synthesized datasets for lifting poses from 2D to 3D and then utilize this to re-target motion from wild settings onto arbitrary avatars. The L3D-Pose dataset can be found at https://soumyaratnadebnath.github.io/L3D-Pose
Soumyaratna Debnath, Harish Katti, Shashikant Verma, Shanmuganathan Raman
ICASSP4
2025 BloomCoreset: Fast Coreset Sampling using Bloom Filters for Fine-Grained Self-Supervised Learning
abstract
The success of deep learning in supervised fine-grained recognition for domain-specific tasks relies heavily on expert annotations. The Open-Set for fine-grained Self-Supervised Learning (SSL) problem aims to enhance performance on downstream tasks by strategically sampling a subset of images (the Core-Set) from a large pool of unlabeled data (the Open-Set). In this paper, we propose a novel method, BloomCoreset, that significantly reduces sampling time from Open-Set while preserving the quality of samples in the coreset. To achieve this, we utilize Bloom filters as an innovative hashing mechanism to store both low- and high-level features of the fine-grained dataset, as captured by Open-CLIP, in a space-efficient manner that enables rapid retrieval of the coreset from the Open-Set. To show the effectiveness of the sampled coreset, we integrate the proposed method into the state-of-the-art fine-grained SSL framework, SimCore [1]. The proposed algorithm drastically outperforms the sampling strategy of the baseline in [1] with a 98.5% reduction in sampling time with a mere 0.83% average trade-off in accuracy calculated across 11 downstream datasets. We have made the code publicly available.
Prajwal Singh, Gautam Vashishtha, Indra Deep Mastan, Shanmuganathan Raman
ICASSP4
2025 Investigating Robustness of Unsupervised Stylegan Image Restoration
abstract
Recently, generative priors have shown significant improvement for unsupervised image restoration. This study explores the incorporation of multiple loss functions that capture various perceptual and structural aspects of image quality. Our proposed method improves robustness across multiple tasks, including denoising, upsampling, inpainting, and deartifacting, by utilizing a comprehensive loss function based on Learned Perceptual Image Patch Similarity(LPIPS), MultiScale Structural Similarity Index Measure Loss(MS-SSIM), Consistency, Feature, and Gradient losses. The experimental results demonstrate marked improvements in accuracy, fidelity, and visual realism in unsupervised image restoration, showcasing the effectiveness of our approach in delivering high-quality results. The experimental results validate the superiority of our approach and offer a promising direction for future advancements in generative-based image restoration methods. Code & Data can be found here https://aamaanakbar.github.io/investigating_rusir/
Indra Deep Mastan, Shanmuganathan Raman
ICIP3
2025 COT-AD: Cotton Analysis Dataset
abstract
This paper presents COT-AD, a comprehensive Dataset designed to enhance cotton crop analysis through computer vision. Comprising over 25,000 images captured throughout the cotton growth cycle, with 5,000 annotated images, COT-AD includes aerial imagery for field-scale detection and segmentation and high-resolution DSLR images documenting key diseases. The annotations cover pest and disease recognition, vegetation, and weed analysis, addressing a critical gap in cotton-specific agricultural datasets. COT-AD supports tasks such as classification, segmentation, image restoration, enhancement, deep generative model-based cotton crop synthesis, and early disease management, advancing data-driven crop management .The COT-AD dataset can be found here: https://aamaanakbar.github.io/COT-AD/.
Mahek Vyas, Soumyaratna Debnath, Chanda Grover Kamra, Jaidev Sanjay Khalane, Reuben Shibu Devanesan, Indra Deep Mastan, Subramanian Sankaranarayanan, Pankaj Khanna, Shanmuganathan Raman
ICIP10
2025 Deep Unsupervised Despeckling With Unbiased Risk Estimation
abstract
Despeckling of Synthetic Aperture Radar (SAR) images has seen significant progress in recent years, largely driven by advancements in deep learning techniques. However, many of these approaches face challenges when applied to new SAR datasets, primarily due to their dependence on ground truth images, which are often unavailable for real-world sensors. In this paper, we address this limitation by extending the concept of unbiased risk estimation in the presence of Gamma-distributed multiplicative speckle. Specifically, we demonstrate that it is possible to train deep denoising networks without relying on ground truth data using our estimator. We introduce a new formulation of the Multiplicative Unbiased Risk Estimator (MURE) and present a computationally efficient Monte Carlo-based method that enables accurate estimation of the modified MURE cost, facilitating effective unsupervised training of deep neural networks from large datasets consisting solely of noisy SAR images. Experimental results on both synthetic datasets and real Sentinel-1 SAR images validate the suitability of our method for real-world applications. Even without ground truth, our method achieves performance that closely matches the Oracle-based denoiser and proves superior to the out-of-domain performance of popular supervised SAR despeckling methods.
Ashutosh Gupta 0008, Chandra Sekhar Seelamantula, Thierry Blu, Nitant Dube, Shanmuganathan Raman
ICIP5
2025 Transformer Augmented Multi-Resolution Hash Encoding in Diffusion Model for 3D Point Cloud Denoising
abstract
Denoising 3D point cloud strives to remove noise from noisy data. Existing methods address the problem by estimating point-wise displacement from the point feature or by learning the distribution of noise. In this paper, we propose to embed the point cloud through a novel multi-resolution hash encoding, and utilize the embedding to learn an optimum transport plan between noisy and corresponding clean point cloud via transformer encoder and a shared-MLP based decoder. The multi-resolution hash encoding uses hierarchical hash-based representations to efficiently capture geometric details at multiple resolutions. Hence, it enables removal of noise while encoding global-to-local structural details. The transformer encoder further improves the model’s ability to learn long-range dependencies and contextual relationships, facilitating improved denoising performance. The optimum transport plan is devised by simulating a denoising diffusion probabilistic model through Schrödinger bridge problem. The proposed method advances state-of-the-art methods through extensive experiments and offers new insights into the synergy between hash encoding, and transformer architectures in the diffusion framework.
Seema Kumari, Utkarsh Mishra, Srimanta Mandal, Shanmuganathan Raman
ICIP4
2025 Darts: Deformable Animation Ready Templates for Clothing Humans
abstract
Accurate 3D modeling of humans and high-fidelity garments is crucial in computer vision and graphics, impacting gaming, virtual, and augmented reality applications. While recent data-driven approaches have progressed in estimating segregated geometries for clothed humans, they often struggle with the seamless integration required for physics-based simulations. We introduce Deformable Animation Ready Templates (DARTs) to address these challenges, which enhance template-based garment reconstruction. Our framework employs a robust feature-line regressor network to establish precise deformation constraints guided by input image characteristics. Additionally, we present a novel differentiable Constrained Rigid Deformation Layer (CRDL) that facilitates effective template deformation while preserving the essential geometry of the garment. Our experiments demonstrate that DARTs can generate templates for physics-based simulation, allowing for seamless garment animations influenced by dynamic environmental factors. With minor adjustments, our templates can accommodate various clothing categories, promoting diversity in animated garment modeling.
Shashikant Verma, Shanmuganathan Raman
ICIP2
2025 GMOT-Mamba: Mamba-Based Model Prediction For Generic Multiple Object Tracking
abstract
We introduce GMOT-Mamba, a novel Mamba-based model prediction framework for Generic Multiple Object Tracking (GMOT) in video sequences. Our approach features a Weighted Feature Pooling (WFP) layer, which processes encoded target states, and an innovative encoder-decoder architecture that leverages Vision-Mamba (ViM) to predict filter weights. We train our model on combinations of large-scale datasets to capture strong priors and discriminative features necessary for generic object tracking. Through extensive experiments and ablation studies, we demonstrate the effectiveness of our approach, showcasing its competitive performance against state-of-the-art GMOT methods while outperforming SOT methods in both accuracy and inference speed. Our findings underscore the potential of Mamba for enhancing model prediction in visual tracking applications.
Shashikant Verma, Nicu Sebe, Shanmuganathan Raman
ICIP3
2025 LIPIDS: Learning-based Illumination Planning In Discretized (Light) Space for Photometric Stereo
abstract
Photometric stereo is a powerful technique for estimating per-pixel surface normals from images under varied il-lumination. Although several methods address photometric stereo with different image (or light) counts ranging from one to two to a hundred, very few focus on learning optimal lighting configuration. Finding an optimal configuration is challenging due to the large number of possible lighting di-rections. Moreover, exhaustive sampling of all possibilities is impractical due to time and resource constraints. Pho-tometric stereo methods have demonstrated promising per-formance on existing datasets, which feature limited light directions sparsely sampled from the light space. There-fore, can we optimally utilize these datasets for illumination planning? In this work, we introduce LIPIDS - Learning-based Illumination Planning In Discretized light Space to achieve minimal and optimal lighting configurations for photometric stereo under arbitrary light distribution. We propose a Light Sampling Network (LSNet) that optimizes the lighting direction for a fixed number of lights by min-imizing the normal loss through a normal regression net-work. The learned light configurations can directly estimate surface normals during inference, even using an off-the-shelf photometric stereo method. Extensive qualitative and quantitative analysis on synthetic and real-world datasets show that photometric stereo under learned lighting config-urations through LIPIDS either surpasses or is nearly com-parable to existing illumination planning methods across different photometric stereo backbones.
Ashish Tiwari 0005, Mihir Sutariya, Shanmuganathan Raman
WACV3
2025 TensoIS: A Step Towards Feed-Forward Tensorial Inverse Subsurface Scattering for Perlin Distributed Heterogeneous Media
abstract
Abstract Estimating scattering parameters of heterogeneous media from images is a severely under‐constrained and challenging problem. Most of the existing approaches model BSSRDF either through an analysis‐by‐synthesis approach, approximating complex path integrals, or using differentiable volume rendering techniques to account for heterogeneity. However, only a few studies have applied learning‐based methods to estimate subsurface scattering parameters, but they assume homogeneous media. Interestingly, no specific distribution is known to us that can explicitly model the heterogeneous scattering parameters in the real world. Notably, procedural noise models such as Perlin and Fractal Perlin noise have been effective in representing intricate heterogeneities of natural, organic, and inorganic surfaces. Leveraging this, we first create HeteroSynth, a synthetic dataset comprising photorealistic images of heterogeneous media whose scattering parameters are modeled using Fractal Perlin noise. Furthermore, we propose Tensorial Inverse Scattering (TensoIS), a learning‐based feed‐forward framework to estimate these Perlin‐distributed heterogeneous scattering parameters from sparse multi‐view image observations. Instead of directly predicting the 3D scattering parameter volume, TensoIS uses learnable low‐rank tensor components to represent the scattering volume. We evaluate TensoIS on unseen heterogeneous variations over shapes from the HeteroSynth test set, smoke and cloud geometries obtained from open‐source realistic volumetric simulations, and some real‐world samples to establish its effectiveness for inverse scattering. Overall, this study is an attempt to explore Perlin noise distribution, given the lack of any such well‐defined distribution in literature, to potentially model real‐world heterogeneous scattering in a feed‐forward manner. Project Page: https://yashbachwana.github.io/TensoIS/
Ashish Tiwari 0005, Satyam Bhardwaj, Yash Bachwana, Parag Sarvoday Sahu, T. M. Feroz Ali, Bhargava Chintalapati, Shanmuganathan Raman
Comput. Graph. Forum7
2025 Structure preserving point cloud completion and classification with coarse-to-fine information
Seema Kumari, Srimanta Mandal, Shanmuganathan Raman
J. Vis. Commun. Image Represent.3
2025 Contrastive Attention-Based Network for Self-Supervised Point Cloud Completion
abstract
Point cloud completion aims to reconstruct complete 3D shapes from partial observations, often requiring multiple views or complete data for training. In this paper, we propose an attention-driven, self-supervised autoencoder network that completes 3D point clouds from a single partial observation. Multi-head self-attention captures robust contextual relationships, while residual connections in the autoencoder enhance geometric feature learning. In addition to this, we incorporate a contrastive learning-based loss, which encourages the network to better distinguish structural patterns even in highly incomplete observations. Experimental results on benchmark datasets demonstrate that the proposed approach achieves state-of-the-art performance in single-view point cloud completion.
Seema Kumari, Preyum Kumar, Srimanta Mandal, Shanmuganathan Raman
IEEE Signal Process. Lett.4
2024 MERLiN: Single-Shot Material Estimation and Relighting for Photometric Stereo
Ashish Tiwari 0005, Satoshi Ikehata, Shanmuganathan Raman
ECCV (12)3
2024 PointGADM: Geometry Acquainted Deep Model for 3D Point Cloud Analysis
Seema Kumari, Samay Kalpesh Patel, Raja Muthalagu, Shanmuganathan Raman
ICPR (30)4
2024 SemFaceEdit: Semantic Face Editing on Generative Radiance Manifolds
Shashikant Verma, Shanmuganathan Raman
ICPR (6)2
2024 Learning Robust Deep Visual Representations from EEG Brain Recordings
abstract
Decoding the human brain has been a hallmark of neuroscientists and Artificial Intelligence researchers alike. Reconstruction of visual images from brain Electroencephalography (EEG) signals has garnered a lot of interest due to its applications in brain-computer interfacing. This study proposes a two-stage method where the first step is to obtain EEG-derived features for robust learning of deep representations and subsequently utilize the learned representation for image generation and classification. We demonstrate the generalizability of our feature extraction pipeline across three different datasets using deep-learning architectures with supervised and contrastive learning methods. We have performed the zero-shot EEG classification task to support the generalizability claim further. We observed that a subject invariant linearly separable visual representation was learned using EEG data alone in an unimodal setting that gives better k-means accuracy as compared to a joint representation learning between EEG and images. Finally, we propose a novel framework to transform unseen images into the EEG space and reconstruct them with approximation, showcasing the potential for image reconstruction from EEG signals. Our proposed image synthesis method from EEG shows 62.9% and 36.13% inception score improvement on the EEGCVPR40 and the Thoughtviz datasets, which is better than state-of-the-art performance in GAN1.
Prajwal Singh, Dwip Dalal, Gautam Vashishtha, Krishna P. Miyapuram, Shanmuganathan Raman
WACV5
2024 GraphFill: Deep Image Inpainting using Graphs
abstract
We present a novel coarser-to-finer approach for deep graphical image inpainting that utilizes GraphFill, a graph neural network-based deep learning framework, and a lightweight generative baseline network. We construct a pyramidal graph for the input-masked image by reducing it into superpixels, each representing a node in the graph. The proposed pyramidal approach facilitates the transfer of global context from coarser to finer pyramid levels, enabling GraphFill to estimate plausible information for unknown node values in the graph. The estimated information is used to fill in the masked region, which a Refine Network then refines. Furthermore, we propose a resolution-robust pyramidal graph construction method, allowing for efficient inpainting of high-resolution images with relatively fewer computations. Our proposed GAN-based network is trained in adversarial settings on Places365 and CelebA-HQ datasets and demonstrates competitive performance compared to existing methods while using fewer learning parameters. We conduct thorough ablation studies to evaluate the effectiveness of each component in the GraphFill Network for improved performance. Our proposed lightweight model for image inpainting is efficient in real-world scenarios, as it can be easily deployed on mobile devices with limited resources.
Shashikant Verma, Roopa Sheshadri, Shanmuganathan Raman
WACV4
2024 Search Me Knot, Render Me Knot: Embedding Search and Differentiable Rendering of Knots in 3D
abstract
Abstract We introduce the problem of knot‐based inverse perceptual art. Given multiple target images and their corresponding viewing configurations, the objective is to find a 3D knot‐based tubular structure whose appearance resembles the target images when viewed from the specified viewing configurations. To solve this problem, we first design a differentiable rendering algorithm for rendering tubular knots embedded in 3D for arbitrary perspective camera configurations. Utilizing this differentiable rendering algorithm, we search over the space of knot configurations to find the ideal knot embedding. We represent the knot embeddings via homeomorphisms of the desired template knot, where the weights of an invertible neural network parametrize the homeomorphisms. Our approach is fully differentiable, making it possible to find the ideal 3D tubular structure for the desired perceptual art using gradient‐based optimization. We propose several loss functions that impose additional physical constraints, enforcing that the tube is free of self‐intersection, lies within a predefined region in space, satisfies the physical bending limits of the tube material, and the material cost is within a specified budget. We demonstrate through results that our knot representation is highly expressive and gives impressive results even for challenging target images in both single‐view and multiple‐view constraints. Through extensive ablation study, we show that each proposed loss function effectively ensures physical realizability. We construct a real‐world 3D‐printed object to demonstrate the practical utility of our approach.
Aalok Gangopadhyay, Paras Gupta, Prajwal Singh, Shanmuganathan Raman
Comput. Graph. Forum5
2023 EEG2IMAGE: Image Reconstruction from EEG Brain Signals
abstract
Reconstructing images using brain signals of imagined visuals may provide an augmented vision to the disabled, leading to the advancement of Brain-Computer Interface (BCI) technology. The recent progress in deep learning has boosted the study area of synthesizing images from brain signals using Generative Adversarial Networks (GAN). In this work, we have proposed a framework for synthesizing the images from the brain activity recorded by an electroencephalogram (EEG) using small-size EEG datasets. This brain activity is recorded from the subject’s head scalp using EEG when they ask to visualize certain classes of Objects and English characters. We use a contrastive learning method in the proposed framework to extract features from EEG signals and synthesize the images from extracted features using conditional GAN. We modify the loss function to train the GAN, which enables it to synthesize 128 × 128 images using a small number of images. Further, we conduct ablation studies and experiments to show the effectiveness of our proposed framework over other state-of-the-art methods using the small EEG dataset.
Prajwal Singh, Pankaj Pandey, Krishna P. Miyapuram, Shanmuganathan Raman
ICASSP4
2023 Single Image LDR to HDR Conversion Using Conditional Diffusion
abstract
Digital imaging aims to replicate realistic scenes, but Low Dynamic Range (LDR) cameras cannot represent the wide dynamic range of real scenes, resulting in under-/overexposed images. This paper presents a deep learning-based approach for recovering intricate details from shadows and highlights while reconstructing High Dynamic Range (HDR) images. We formulate the problem as an image-to-image (I2I) translation task and propose a conditional Denoising Diffusion Probabilistic Model (DDPM) based framework using classifier-free guidance. We incorporate a deep CNN-based autoencoder in our proposed framework to enhance the quality of the latent representation of the input LDR image used for conditioning. Moreover, we introduce a new loss function for LDR-HDR translation tasks, termed Exposure Loss. This loss helps direct gradients in the opposite direction of the saturation, further improving the results’ quality. By conducting comprehensive quantitative and qualitative experiments, we have effectively demonstrated the proficiency of our proposed method. The results indicate that a simple conditional diffusion-based method can replace the complex camera pipeline-based architectures.
Dwip Dalal, Gautam Vashishtha, Prajwal Singh, Shanmuganathan Raman
ICIP4
2022 DeepPS2: Revisiting Photometric Stereo Using Two Differently Illuminated Images
Ashish Tiwari 0005, Shanmuganathan Raman
ECCV (7)2
2022 Exploring Deeper Graph Convolutions for Semi-Supervised Node Classification
abstract
Graph convolutional networks (GCNs) have achieved impressive performance in learning from graph-structured data. Although GCN and its variants have shown promising results, they continue to remain shallow as their performance drops with an increasing number of layers - a problem popularly known as oversmoothing. This work introduces a simple yet effective idea of feature gating over graph convolution layers to facilitate deeper graph neural networks and address oversmoothing. The proposed feature gating is easy to implement without changing the underlying network architecture and is broadly applicable to GCN and almost any of its variants. Further, we demonstrate the use of feature gating in assigning importance to node features and the nodes for the node classification task. Quantitative analysis on real-world datasets shows that feature gating paves the way for constructing deeper GCNs.
Ashish Tiwari 0005, Richeek Das, Shanmuganathan Raman
ICASSP3
2022 LERPS: Lighting Estimation and Relighting for Photometric Stereo
abstract
Photometric stereo is a method to obtain surface normals of an object using its images captured under varying illumination directions. The existing deep learning-based methods require multiple images of an object captured using complex image acquisition systems. In this work, we propose a deep learning framework to perform three tasks jointly: (i) lighting estimation, (ii) image relighting, and (iii) surface normal estimation, all from a single input image of an object with non-Lambertian surface and general reflectance. The network explicitly segregates global geometric features and local lighting-specific features of the object from a single image. The local features resemble attached shadows, shadings, and specular highlights, providing valuable lighting estimation and relighting cues. The global features capture the lighting-independent geometric attributes that effectively guide the surface normal estimation. The joint training transfers valuable insights to achieve significant improvements across all three tasks. We show that the proposed single-image-based relighting framework outperforms several existing photometric stereo methods which require multiple images of a static object.
Ashish Tiwari 0005, Shanmuganathan Raman
ICASSP2
2022 DMD-Net: Deep Mesh Denoising Network
abstract
We present Deep Mesh Denoising Network (DMD-Net), an end-to-end deep learning framework, for solving the mesh denoising problem. DMD-Net consists of a Graph Convolutional Neural Network in which aggregation is performed in both the primal as well as the dual graph. This is realized in the form of an asymmetric two-stream network, which contains a primal-dual fusion block that enables communication between the primal-stream and the dual-stream. We develop a Feature Guided Transformer (FGT) paradigm, which consists of a feature extractor, a transformer, and a denoiser. The feature extractor estimates the local features, that guide the transformer to compute a transformation, which is applied to the noisy input mesh to obtain a useful intermediate representation. This is further processed by the denoiser to obtain the denoised mesh. Our network is trained on a large scale dataset of 3D objects. We perform exhaustive ablation studies to demonstrate that each component in our network is essential for obtaining the best performance. We show that our method obtains competitive or better results when compared with the state-of-the-art mesh denoising algorithms. We demonstrate that our method is robust to various kinds of noise. We observe that even in the presence of extremely high noise, our method achieves excellent performance.
Aalok Gangopadhyay, Shashikant Verma, Shanmuganathan Raman
ICPR3
2022 DeepHS-HDRVideo: Deep High Speed High Dynamic Range Video Reconstruction
abstract
Due to hardware constraints, standard off-the-shelf digital cameras suffers from low dynamic range (LDR) and low frame per second (FPS) outputs. Previous works in high dynamic range (HDR) video reconstruction uses sequence of alternating exposure LDR frames as input, and align the neighbouring frames using optical flow based networks. However, these methods often result in motion artifacts in challenging situations. This is because, the alternate exposure frames have to be exposure matched in order to apply alignment using optical flow. Hence, over-saturation and noise in the LDR frames results in inaccurate alignment. To this end, we propose to align the input LDR frames using a pre-trained video frame interpolation network. This results in better alignment of LDR frames, since we circumvent the error-prone exposure matching step, and directly generate intermediate missing frames from the same exposure inputs. Furthermore, it allows us to generate high FPS HDR videos by recursively interpolating the intermediate frames. Through this work, we propose to use video frame interpolation for HDR video reconstruction, and present the first method to generate high FPS HDR videos. Experimental results demonstrate the efficacy of the proposed framework against optical flow based alignment methods, with an absolute improvement of 2.4 PSNR value on standard HDR video datasets [1], [2] and further benchmark our method for high FPS HDR video generation.
Zeeshan Khan, Parth Shettiwar, Mukul Khanna, Shanmuganathan Raman
ICPR4
2022 LS-HDIB: A Large Scale Handwritten Document Image Binarization Dataset
abstract
Handwritten document image binarization is challenging due to high variability in the written content and complex background attributes such as page style, paper quality, stains, shadow gradients, and non-uniform illumination. While the traditional thresholding methods do not effectively generalize on such challenging real-world scenarios, deep learning-based methods have performed relatively well when provided with sufficient training data. However, the existing datasets are limited in size and diversity. This work proposes LS-HDIB - a large-scale handwritten document image binarization dataset containing over a million document images that span numerous real-world scenarios. Additionally, we introduce a novel technique that uses a combination of adaptive thresholding and seamless cloning methods to create the dataset with accurate ground truths. Through an extensive quantitative and qualitative evaluation over eight different deep learning based models, we demonstrate the enhancement in the performance of these models when trained on the LS-HDIB dataset and tested on unseen images.
Kaustubh Sadekar, Ashish Tiwari 0005, Prajwal Singh, Shanmuganathan Raman
ICPR4
2022 Deep Appearance Consistent Human Pose Transfer
abstract
The fidelity of a pose transfer system depends on its ability to generate realistic images of a person under novel poses while preserving the desired human attributes (like face, hairstyle, and clothes). However, the visual fidelity is often compromised as the existing methods fail to extract rich appearance and pose features since they propagate the pose and the appearance information through the same pathway. Also, the repeated downsampling in these pathways leads to the loss of finer details, thus producing blurry results. Further, these methods use vanilla convolution that treats all the pixels as important and fail to focus primarily on significant regions needed for the desired transformation. This work proposes an appearance-consistent human pose transfer framework that progressively transforms the person in the source image to the desired target pose using the information from three pathways: an image pathway, a pose pathway, and an appearance pathway. We propose the use of gated convolution to dynamically extract features relevant for generating the transformed image. The appearance pathway generates an appearance code to produce an image consistent in appearance with that of the source image. We establish the efficacy of the proposed framework through an extensive set of experiments on DeepFashion, Market-1501, and the Action Class dataset. We also generate coherent action sequences through a given set of desired poses from the action class dataset that contains humans in three actions: golf, yoga/workouts, and tennis.
Ashish Tiwari 0005, Zeeshan Khan, Aditya Vora, Manjuprakash Rama Rao, Shanmuganathan Raman
ICPR5
2022 RGL-NET: A Recurrent Graph Learning framework for Progressive Part Assembly
abstract
Autonomous assembly of objects is an essential task in robotics and 3D computer vision. It has been studied extensively in robotics as a problem of motion planning, actuator control and obstacle avoidance. However, the task of developing a generalized framework for assembly robust to structural variants remains relatively unexplored. In this work, we tackle this problem using a recurrent graph learning framework considering inter-part relations and the progressive update of the part pose. Our network can learn more plausible predictions of shape structure by accounting for priorly assembled parts. Compared to the current state-of-the-art, our network yields up to 10% improvement in part accuracy and up to 15% improvement in connectivity accuracy on the PartNet [23] dataset. Moreover, our resulting latent space facilitates exciting applications such as shape recovery from the point-cloud components. We conduct extensive experiments to justify our design choices and demonstrate the effectiveness of the proposed framework.
Abhinav Narayan Harish, Rajendra Nagar, Shanmuganathan Raman
WACV3
2022 Shadow Art Revisited: A Differentiable Rendering Based Approach
abstract
While recent learning-based methods have been observed to be superior for several vision-related applications, their potential in generating artistic effects has not been explored much. One such exciting application is Shadow Art - a unique form of sculptural art that produces artistic effects through 2D shadows cast by a 3D sculpture. In this work, we revisit shadow art using differentiable rendering-based optimization frameworks to obtain the 3D sculpture from a set of shadow (binary) images and their corresponding projection information. Specifically, we discuss shape optimization through voxel as well as mesh-based differentiable renderers. Our choice of using differentiable rendering for generating shadow art sculptures can be attributed to its ability to learn the underlying 3D geometry solely from image data, thus reducing the dependence on 3D ground truth. The qualitative and quantitative results demonstrate the potential of the proposed framework in generating complex 3D sculptures that transcend the ones seen in contemporary art pieces using just a set of shadow images as input. Further, we demonstrate the generation of 3D sculptures to cast shadows of faces, animated movie characters, and the applicability of the proposed framework to sketch-based 3D reconstruction of the underlying shapes.
Kaustubh Sadekar, Ashish Tiwari 0005, Shanmuganathan Raman
WACV3
2022 Depthwise Spatio-Temporal STFT Convolutional Neural Networks for Human Action Recognition
abstract
Conventional 3D convolutional neural networks (CNNs) are computationally expensive, memory intensive, prone to overfitting, and most importantly, there is a need to improve their feature learning capabilities. To address these issues, we propose spatio-temporal short term Fourier transform (STFT) blocks, a new class of convolutional blocks that can serve as an alternative to the 3D convolutional layer and its variants in 3D CNNs. An STFT block consists of non-trainable convolution layers that capture spatially and/or temporally local Fourier information using a STFT kernel at multiple low frequency points, followed by a set of trainable linear weights for learning channel correlations. The STFT blocks significantly reduce the space-time complexity in 3D CNNs. In general, they use 3.5 to 4.5 times less parameters and 1.5 to 1.8 times less computational costs when compared to the state-of-the-art methods. Furthermore, their feature learning capabilities are significantly better than the conventional 3D convolutional layer and its variants. Our extensive evaluation on seven action recognition datasets, including Something-something v1 and v2, Jester, Diving-48, Kinetics-400, UCF 101, and HMDB 51, demonstrate that STFT blocks based 3D CNNs achieve on par or even better performance compared to the state-of-the-art methods.
Sudhakar Kumawat, Manisha Verma, Yuta Nakashima, Shanmuganathan Raman
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 3d Point Cloud Completion Using Stacked Auto-Encoder For Structure Preservation
abstract
3D point cloud completion problem deals with completing the shape from partial points. The problem finds its application in many vision-related applications. Here, structure plays an important role. Most of the existing approaches either do not consider structural information or consider structure at the decoder only. For maintaining the structure, it is also necessary to maintain the position of the available 3D points. However, most of the approaches lack the aspect of maintaining the available structural position. In this paper, we propose to employ stacked auto-encoder in conjunction a with shared Multi-Layer Perceptron (MLP). MLP converts each 3D point into a feature vector and the stacked auto-encoder helps in maintaining the available structural position of the input points. Further, it explores the redundancy present in the feature vector. It aids to incorporate coarse to fine scale information that further helps in better shape representation. The embedded feature is finally decoded by a structural preserving decoder. Both the encoding and the decoding operations of our method take care of preserving the structure of the available shape information. The experimental results demonstrate the structure preserving capability of our network as compared to the state-of-the-art methods.
Seema Kumari, Shanmuganathan Raman
ICIP2
2021 DeepCFL: Deep Contextual Features Learning from a Single Image
abstract
Recently, there is a vast interest in developing image feature learning methods that are independent of the training data, such as deep image prior [35], InGAN [28], [29], SinGAN [27], and DCIL [8]. These methods perform various tasks, such as image restoration, image editing, and image synthesis. In this work, we proposed a new training data-independent framework, called Deep Contextual Features Learning (DeepCFL), to perform image synthesis and image restoration based on the semantics of the input image. The contextual features are simply the high dimensional vectors representing the semantics of the given image. DeepCFL is a single image GAN framework that learns the distribution of the context vectors from the input image. We show the performance of contextual learning in various challenging scenarios: outpainting, inpainting, and restoration of randomly removed pixels. DeepCFL is applicable when the input source image and the generated target image are not aligned. We illustrate image synthesis using DeepCFL for the task of image resizing.
Indra Deep Mastan, Shanmuganathan Raman
WACV2
2020 Depthwise-STFT Based Separable Convolutional Neural Networks
abstract
In this paper, we propose a new convolutional layer called Depthwise-STFT Separable layer that can serve as an alternative to the standard depthwise separable convolutional layer. The construction of the proposed layer is inspired by the fact that the Fourier coefficients can accurately represent important features such as edges in an image. It utilizes the Fourier coefficients computed (channelwise) in the 2D local neighborhood (e.g., 3 × 3) of each position of the input map to obtain the feature maps. The Fourier coefficients are computed using 2D Short Term Fourier Transform (STFT) at multiple fixed low frequency points in the 2D local neighborhood at each position. These feature maps at different frequency points are then linearly combined using trainable pointwise (1 × 1) convolutions. We show that the proposed layer outperforms the standard depthwise separable layer based models on the CIFAR-10 and CIFAR-100 image classification datasets with reduced space-time complexity.
Sudhakar Kumawat, Shanmuganathan Raman
ICASSP2
2020 Simultaneous Detection and Removal of Dynamic Objects in Multi-view Images
abstract
Consider a set of images of a scene consisting of moving objects captured using a hand-held camera. In this work, we propose an algorithm which takes this set of multi-view images as input, detects the dynamic objects present in the scene, and replaces them with the static regions which are being occluded by them. The proposed algorithm scans the reference image in the row-major order at the pixel level and classifies each pixel as static or dynamic. During the scan, when a pixel is classified as dynamic, the proposed algorithm replaces that pixel value with the corresponding pixel value of the static region which is being occluded by that dynamic region. We show that we achieve artifact-free removal of dynamic objects in multi-view images of several real-world scenes. To the best of our knowledge, we propose the first method which simultaneously detects and removes the dynamic objects present in multi-view images.
Gagan Kanojia, Shanmuganathan Raman
WACV2
2020 DCIL: Deep Contextual Internal Learning for Image Restoration and Image Retargeting
abstract
Recently, there is a vast interest in developing unsupervised methods that are independent of the feature learning from the training data, e.g., deep image prior [26], zero-shot learning [23], and internal learning [21], [22]. These methods are based on the common goal of maxi-mizing the quality of image features learned from a single image despite inherent technical diversity. In this work, we bridge the gap between the various unsupervised approaches above and propose a general framework for image restoration and image retargeting. We use contextual feature learning and internal learning to improvise the structure similarity between the source and the target images. We perform image resizing application in the following setups: classical image resizing using super-resolution, a challenging image resizing where the low-resolution image contains noise, and content-aware image resizing using image retar-geting. We also compare our framework with relevant state-of-the-art methods.
Indra Deep Mastan, Shanmuganathan Raman
WACV2
2020 3DSymm: Robust and Accurate 3D Reflection Symmetry Detection
Rajendra Nagar, Shanmuganathan Raman
Pattern Recognit.2
2019 LP-3DCNN: Unveiling Local Phase in 3D Convolutional Neural Networks
abstract
Traditional 3D Convolutional Neural Networks (CNNs) are computationally expensive, memory intensive, prone to overfit, and most importantly, there is a need to improve their feature learning capabilities. To address these issues, we propose Rectified Local Phase Volume (ReLPV) block, an efficient alternative to the standard 3D convolutional layer. The ReLPV block extracts the phase in a 3D local neighborhood (e.g., 3 × 3 × 3) of each position of the input map to obtain the feature maps. The phase is extracted by computing 3D Short Term Fourier Transform (STFT) at multiple fixed low frequency points in the 3D local neighborhood of each position. These feature maps at different frequency points are then linearly combined after passing them through an activation function. The ReLPV block provides significant parameter savings of at least, 33to 133times compared to the standard 3D convolutional layer with the filter sizes 3 × 3 × 3 to 13 × 13 × 13, respectively. We show that the feature learning capabilities of the ReLPV block are significantly better than the standard 3D convolutional layer. Furthermore, it produces consistently better results across different 3D data representations. We achieve state-of-the-art accuracy on the volumetric ModelNet10 and ModelNet40 datasets while utilizing only 11% parameters of the current state-of-the-art. We also improve the state-of-the-art on the UCF-101 split-1 action recognition dataset by 5.68% (when trained from scratch) while using only 15% of the parameters of the state-of-the-art.
Sudhakar Kumawat, Shanmuganathan Raman
CVPR2
2019 Local Phase U-net for Fundus Image Segmentation
abstract
In this paper, we propose Rectified Local Phase Unit (ReLPU), which is an efficient and trainable convolutional layer that utilizes phase information computed locally in a window for every pixel location of the input image. The ReLPU layer is based on applying the Rectified Linear Unit (ReLU) activation function on the local phase information extracted by computing the local Fourier transform of the input image at multiple low frequency points. The ReLPU layer, when used at the top of the segmentation network U-Net, is observed to improve the performance of the baseline U-Net model. We demonstrate this using the task of segmenting blood vessels in fundus images of two standard datasets, DRIVE and STARE, achieving state-of-the-art results. An important feature of the ReLPU layer is that it is trainable which allows it to choose the best frequency points for computing local Fourier transform and to selectively give more weight to them during training.
Sudhakar Kumawat, Shanmuganathan Raman
ICASSP2
2019 Reflection Symmetry Detection by Embedding Symmetry in a Graph
abstract
Reflection symmetry is ubiquitous in nature and plays an important role in object detection and recognition tasks. Most of the existing methods for symmetry detection extract and describe each keypoint using a descriptor and a mirrored descriptor. Two keypoints are said to be mirror symmetric key-points if the original descriptor of one keypoint and the mirrored descriptor of the other keypoint are similar. However, these methods suffer from the following issue. The background pixels around the mirror symmetric pixels lying on the boundary of an object can be different. Therefore, their descriptors can be different. However, the boundary of a symmetric object is a major component of global reflection symmetry. We exploit the estimated boundary of the object and describe a boundary pixel using only the estimated normal of the boundary segment around the pixel. We embed the symmetry axes in a graph as cliques to robustly detect the symmetry axes. We show that this approach achieves state-of-the-art results in a standard dataset.
Rajendra Nagar, Shanmuganathan Raman
ICASSP2
2019 Accelerated seam carving for image retargeting
abstract
Display of images on different display devices having varied size and aspect ratio requires one to resize them. Many attempts have been made to perform content‐aware image retargeting while generating an image compatible with a target display size. Seam carving is one of the image retargeting operators which alters the size of an image by removing least energy pixels. However, it requires high computational time in order to perform retargeting. In this study, the authors accelerate the naive seam carving process by removal or insertion of multiple pixel wide batch seam in a single iteration rather than a single pixel wide seam. Along with the energy of pixels to be removed, inserted energy after the removal of a batch seam is also minimised in order to prevent the inclusion of false edges. The width of a batch seam is a critical factor which is made adaptive during the retargeting process to preserve the energy of an image. They have shown a significant decrease in computational time with the increase in the width of a batch seam. They have compared the proposed technique with other state‐of‐the‐art image retargeting operators using different quality assessment metrics and visual results.
Diptiben Patel, Shanmuganathan Raman
IET Image Process.2
2019 DeepImSeq: Deep image sequencing for unsynchronized cameras
Gagan Kanojia, Shanmuganathan Raman
Pattern Recognit. Lett.2
2019 Reflection symmetry aware image retargeting
Diptiben Patel, Rajendra Nagar, Shanmuganathan Raman
Pattern Recognit. Lett.3
2019 Object occlusion guided stereo image retargeting
Diptiben Patel, Shanmuganathan Raman
Pattern Recognit. Lett.2
2019 Patch-based detection of dynamic objects in CrowdCam images
Gagan Kanojia, Shanmuganathan Raman
Vis. Comput.2
2018 Fast and Accurate Intrinsic Symmetry Detection
Rajendra Nagar, Shanmuganathan Raman
ECCV (1)2
2018 Hashing in the zero shot framework with domain adaptation
Shubham Pachori, Ameya Deshpande, Shanmuganathan Raman
Neurocomputing3
2018 Iterative spectral clustering for unsupervised object localization
Aditya Vora, Shanmuganathan Raman
Pattern Recognit. Lett.2
2017 No-reference quality assessment of tone mapped High Dynamic Range (HDR) images using transfer learning
abstract
We present a transfer learning framework for no-reference image quality assessment (NRIQA) of tonemapped High Dynamic Range (HDR) images. This work is motivated by the observation that quality assessment databases in general, and HDR image databases in particular are “small” relative to the typical requirements for training deep neural networks. Transfer learning based approaches have been successful in such scenarios where learning from a related but larger database is transferred to the smaller database. Specifically, we propose a framework where the successful AlexNet is used to extract image features. This is followed by the application of Principal Component Analysis (PCA) to reduce the dimensionality of the feature vector (from 4096 to 400), given the small database size. A linear regression model is then fit to Mean Opinion Scores (MOS) using L2 regularization to prevent overfitting. We demonstrate state-of-the-art performance of the proposed approach on the ESPL-LIVE database.
Abhinau Kumar Venkataramanan, Shashank Gupta 0001, Sai Sheetal Chandra, Shanmuganathan Raman, Sumohana S. Channappayya
QoMEX4
2017 Postcapture Focusing Using Regression Forest
abstract
A photograph of the same scene can look different when captured using different camera settings. In this letter, we propose a novel technique to obtain the complete postcapture control over the focus and aperture settings of a traditional camera by acquiring small number of images. In this letter, we tackle the problem of deciding which focus-aperture (F-A) combinations should be used to capture the input images, by solving it as a center selection problem. The images captured with the selected settings are then used to reconstruct the images for all possible F-A settings of a traditional camera. For the reconstruction of the images, we have used random regression forest. We show that the proposed approach provides an effective alternative for postcapture control in photography.
Gagan Kanojia, Shanmuganathan Raman
IEEE Signal Process. Lett.2
2017 Reflection Symmetry Axes Detection Using Multiple Model Fitting
abstract
We propose an energy minimization approach to detect multiple reflection symmetry axes present in a given image representing fronto-parallel view of a scene. We perform local feature matching to detect the pairs of mirror symmetric points, and in order to formulate an energy function, we use the geometric characteristics of the symmetry axis. That is, it passes through the midpoint of line segment joining the two mirror symmetric points and is perpendicular to the vector joining two mirror symmetric points. We propose a novel k-symmetry clustering algorithm to minimize this energy function in order to efficiently find all the symmetry axes present in the given image. We evaluate the proposed method on the standard datasets and show that we get comparable and better results than that of the state-of-the-art reflection symmetry detection methods.
Rajendra Nagar, Shanmuganathan Raman
IEEE Signal Process. Lett.2
2016 Revealing Hidden 3-D Reflection Symmetry
abstract
Reflection symmetry is present in most of the man-made or naturally formed objects. In computer vision, real-world scenes are represented by dense 3-D models or by 2-D projections, such as images captured by cameras. Most of the existing methods either detect reflection symmetry from dense 3-D models or 2-D projections. However, generating a dense 3-D model is a computationally expensive process and reflection symmetry may not be evident in any of the 2-D views obtained through projections. In this letter, we propose an energy minimizationbased approach to detect the reflection symmetry present in the object from its multiple 2-D projections captured from different viewpoints and the sparse 3-D model obtained using these projections. The proposed approach only estimates the sparse 3-D model and utilizes content of the images in terms of local scale invariant features. The energy minimization problem reduces to the problem of finding the eigenvector corresponding to the smallest eigenvalue of a small matrix, thereby leading to reduction in computations.
Rajendra Nagar, Shanmuganathan Raman
IEEE Signal Process. Lett.2
2016 Robust PCA-based solution to image composition using augmented Lagrange multiplier (ALM)
Adit Bhardwaj, Shanmuganathan Raman
Vis. Comput.2
2011 Reconstruction of high contrast images for dynamic scenes
Shanmuganathan Raman, Subhasis Chaudhuri
Vis. Comput.1
2009 Poisson compositing
abstract
Most of the real world scenes have a very high dynamic range. However the common capture and display devices can handle only a limited dynamic range. General approach to solve this problem is to use multi-exposure images and composite them in the irradiance domain to get a High Dynamic Range (HDR) image [Reinhard et al. 2005]. The generated image will be able to represent the real world scene faithfully. However, it needs to be tone-mapped to a Low Dynamic Range (LDR) image for visualization in common displays and printers. Generation of the high-quality LDR image of the scene directly from multi-exposure images even in the absence of any knowledge of camera response function and the exposure settings of the camera is of interest to graphics community. We propose a gradient domain compositing technique to solve the above problem and call it Poisson Compositing. We compare the proposed methodology with similar existing techniques and show that the proposed method is very fast and accurate.
Shanmuganathan Raman, Subhasis Chaudhuri
SIGGRAPH ASIA Sketches1
2007 A Matte-less, Variational Approach to Automatic Scene Compositing
abstract
In this paper, we consider the problem of compositing a scene from multiple images. Multiple images, for example, can be obtained by varying the exposure of the camera, by changing the object at focus, or by simply sampling a video sequence at arbitrary time instants. We develop this problem in an optimization framework and then adopt a variational approach to derive a generalized algorithm which will be able to solve diverse applications depending on the nature of the input images. Our approach has distinct advantages over the existing digital compositing techniques, such as alpha matting and alpha blending, which require an explicit preparation of the matte while there is no such requirement in the proposed technique. We demonstrate the usefulness of our approach through results from diverse applications in computer vision.
Shanmuganathan Raman, Subhasis Chaudhuri
ICCV1