Ankit Dhiman

dblp:227/7222 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
8since 2021 · last 2026
0000-0002-9451-2052ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2026 UniC-Lift: Unified 3D Instance Segmentation via Contrastive Learning
abstract
3D Gaussian Splatting (3DGS) and Neural Radiance Fields (NeRF) have advanced novel-view synthesis. Recent methods extend multi-view 2D segmentation to 3D, enabling instance/semantic segmentation for better scene understanding. A key challenge is the inconsistency of 2D instance labels across views, leading to poor 3D predictions. Existing methods use a two-stage approach in which some rely on contrastive learning with hyperparameter-sensitive clustering, while others preprocess labels for consistency. We propose a unified framework that merges these steps, reducing training time and improving performance by introducing a learnable feature embedding for segmentation in Gaussian primitives. This embedding is then efficiently decoded into instance labels through a novel "Embedding-to-Label" process, effectively integrating the optimization. While this unified framework offers substantial benefits, we observed artifacts at the object boundaries. To address the object boundary issues, we propose hard-mining samples along these boundaries. However, directly applying hard mining to the feature embeddings proved unstable. Therefore, we apply a linear layer to the rasterized feature embeddings before calculating the triplet loss, which stabilizes training and significantly improves performance. Our method outperforms baselines qualitatively and quantitatively on the ScanNet, Replica3D, and Messy-Rooms datasets.
Ankit Dhiman, R. Srinath, Jaswanth Reddy, Lokesh R. Boregowda, Venkatesh Babu Radhakrishnan
AAAI1
2025 Reflecting Reality: Enabling Diffusion Models to Produce Faithful Mirror Reflections
abstract
We tackle the problem of generating highly realistic and plausible mirror reflections using diffusion-based generative models. We formulate this problem as an image inpainting task, allowing for more user control over the placement of mirrors during the generation process. To enable this, we create SynMirror, a large-scale dataset of diverse synthetic scenes with objects placed in front of mirrors. SynMirror contains around 198K samples rendered from 66K unique 3D objects, along with their associated depth maps, normal maps and instance-wise segmentation masks, to capture relevant geometric properties of the scene. Using this dataset, we propose a novel depth-conditioned inpainting method called MirrorFusion, which generates high-quality geometrically consistent and photo-realistic mirror reflections given an input image and a mask depicting the mirror region. MirrorFusion outperforms state-of-the-art methods on SynMirror, as demonstrated by extensive quantitative and qualitative analysis. To the best of our knowledge, we are the first to successfully tackle the challenging problem of generating controlled and faithful mirror reflections of an object in a scene using diffusion based models. Syn-Mirror and MirrorFusion open up new avenues for image editing and augmented reality applications for practitioners and researchers alike. The project page is available at: https://val.cds.iisc.ac.in/reflecting-reality.github.io/.
Ankit Dhiman, Manan Shah, Rishubh Parihar, Yash Bhalgat, Lokesh R. Boregowda, Venkatesh Babu Radhakrishnan
3DV1
2025 MirrorVerse: Pushing Diffusion Models to Realistically Reflect the World
abstract
Diffusion models have become central to various image editing tasks, yet they often fail to fully adhere to physical laws, particularly with effects like shadows, reflections, and occlusions. In this work, we address the challenge of generating photorealistic mirror reflections using diffusion-based generative models. Despite extensive training data, existing diffusion models frequently overlook the nuanced details crucial to authentic mirror reflections. Recent approaches have attempted to resolve this by creating synthetic datasets and framing reflection generation as an in-painting task; however, they struggle to generalize across different object orientations and positions relative to the mirror. Our method overcomes these limitations by introducing key augmentations into the synthetic data pipeline: (1) random object positioning, (2) randomized rotations, and (3) grounding of objects, significantly enhancing generalization across poses and placements. To further address spatial relationships and occlusions in scenes with multiple objects, we implement a strategy to pair objects during dataset generation, resulting in a dataset robust enough to handle these complex scenarios. Achieving generalization to real-world scenes remains a challenge, so we introduce a three-stage training curriculum to develop the MirrorFusion 2.0 model to improve real-world performance. We provide extensive qualitative and quantitative evaluations to support our approach. The project page is available at: https://mirror-verse.github.io/.
Ankit Dhiman, Manan Shah, Venkatesh Babu Radhakrishnan
CVPR1
2025 ChromaDistill: Colorizing Monochrome Radiance Fields with Knowledge Distillation
abstract
Colorization is a well-explored problem in the domains of image and video processing. However, extending colorization to 3D scenes presents significant challenges. Re-cent Neural Radiance Field (NeRF) and Gaussian-Splatting (3DGS) methods enable high-quality novel-view synthe-sis for multi-view images. However, the question arises: How can we colorize these 3D representations? This work presents a method for synthesizing colorized novel views from input grayscale multi-view images. Using image or video colorization methods to colorize novel views from these 3D representations naively will yield output with se-vere inconsistencies. We introduce a novel method to use powerful image colorization models for colorizing 3D representations. We propose a distillation-based method that transfers color from these networks trained on natural images to the target 3D representation. Notably, this strat-egy does not add any additional weights or computational overhead to the original representation during inference. Extensive experiments demonstrate that our method produces high-quality colorized views for indoor and outdoor scenes, showcasing significant cross-view consistency advantages over baseline approaches. Our method is agnos-tic to the underlying 3D representation and easily gener-alizable to NeRF and 3DGS methods. Further, we vali-date the efficacy of our approach in several diverse applications: 1.) Infra-Red (IR) multi-view images and 2.) Legacy grayscale multi-view image sequences. Project Webpage: https://val.cds.iisc.ac.in/chroma-distill.github.io/
Ankit Dhiman, R. Srinath, Srinjay Sarkar, Lokesh R. Boregowda, Venkatesh Babu Radhakrishnan
WACV1
2025 Instructive3D: Editing Large Reconstruction Models with Text Instructions
abstract
Transformer based methods have enabled users to create, modify, and comprehend text and image data. Recently proposed Large Reconstruction Models (LRMs) further extend this by providing the ability to generate high-quality 3D models with the help of a single object image. These models, however, lack the ability to manipulate or edit the finer details, such as adding standard design patterns or changing the color and reflectance of the generated objects, thus lacking fine-grained control that may be very helpful in domains such as augmented reality, animation and gaming. Naively training LRMs for this purpose would require generating precisely edited images and 3D object pairs, which is computationally expensive. In this paper, we propose Instructive3D, a novel LRM based model that integrates generation and fine-grained editing, through user text prompts, of 3D objects into a single model. We accomplish this by adding an adapter that performs a diffusion process conditioned on a text prompt specifying edits in the triplane latent space representation of 3D object models. Our method does not require the generation of edited 3D objects. Additionally, Instructive3D allows us to perform geometrically consistent modifications, as the edits done through user-defined text prompts are applied to the triplane latent representation thus enhancing the versatility and precision of 3D objects generated. We compare the objects generated by Instructive3D and a baseline that first generates the 3D object meshes using a standard LRM model and then edits these 3D objects using text prompts when images are provided from the Objaverse LVIS dataset. We find that Instructive3D produces qualitatively superior 3D objects with the properties specified by the edit prompts.
Kunal Kathare, Ankit Dhiman, Vikas K. Gowda, Siddharth Aravindan, Shubham Monga, Basavaraja S. Vandrotti, Lokesh R. Boregowda
WACV2
2023 Strata-NeRF : Neural Radiance Fields for Stratified Scenes
abstract
Neural Radiance Field (NeRF) approaches learn the underlying 3D representation of a scene and generate photo-realistic novel views with high fidelity. However, most proposed settings concentrate on modelling a single object or a single level of a scene. However, in the real world, we may capture a scene at multiple levels, resulting in a layered capture. For example, tourists usually capture a monument’s exterior structure before capturing the inner structure. Modelling such scenes in 3D with seamless switching between levels can drastically improve immersive experiences. However, most existing techniques struggle in modelling such scenes. We propose Strata-NeRF, a single neural radiance field that implicitly captures a scene with multiple levels. Strata-NeRF achieves this by conditioning the NeRFs on Vector Quantized (VQ) latent representations which allow sudden changes in scene structure. We evaluate the effectiveness of our approach in multi-layered synthetic dataset comprising diverse scenes and then further validate its generalization on the real-world RealEstate10K dataset. We find that Strata-NeRF effectively captures stratified scenes, minimizes artifacts, and synthesizes high-fidelity views compared to existing approaches. https://ankitatiisc.github.io/Strata-NeRF/
Ankit Dhiman, R. Srinath, Harsh Rangwani, Rishubh Parihar, Lokesh R. Boregowda, Srinath Sridhar 0002, Venkatesh Babu Radhakrishnan
ICCV1
2023 NTrans-Net: A Multi-Scale Neutrosophic-Uncertainty Guided Transformer Network for Indoor Depth Completion
abstract
Dense depth maps are important constituents in a variety of tasks and have wide ranging applications. However, depth maps captured by indoor depth sensors have an extensive range of missing depth values and are also sparse in nature. Predicting dense depth from sparse input has been widely studied and is solved either as a regression or classification problem. In this paper, we propose a novel representation, termed Unified Ordinal Vectors, to realise the combined advantages of regression and classification methods. To disinter the potential of this representation, we propose NTrans-Net, a novel multi-scale network that can extract hierarchical and complementary information with neutrosophic indeterminacy feature handling. We also propose a dual encoder-decoder transformer structure to handle these neutrosophic domain features with guided attention to better capture inter-modal dependencies for superior depth completion performance. NTrans-Net is designed to be flexible enough to adapt to the dynamic nature of spatial input contexts and be robust to sensor-dependent distributions. We conduct extensive experiments on NYUv2 and ToF18K datasets to demonstrate the superiority of the proposed method in multiple settings, especially in realistic indoor environments as captured by commodity depth sensors.
Akshat Ramachandran, Ankit Dhiman, Basavaraja S. Vandrotti
ICIP2
2022 Everything is There in Latent Space: Attribute Editing and Attribute Style Manipulation by StyleGAN Latent Space Exploration
abstract
Unconstrained Image generation with high realism is now possible using recent Generative Adversarial Networks (GANs). However, it is quite challenging to generate images with a given set of attributes. Recent methods use style-based GAN models to perform image editing by leveraging the semantic hierarchy present in the layers of the generator. We present Few-shot Latent-based Attribute Manipulation and Editing (FLAME), a simple yet effective framework to perform highly controlled image editing by latent space manipulation. Specifically, we estimate linear directions in the latent space (of a pre-trained StyleGAN) that controls semantic attributes in the generated image. In contrast to previous methods that either rely on large-scale attribute labeled datasets or attribute classifiers, FLAME uses minimal supervision of a few curated image pairs to estimate disentangled edit directions. FLAME can perform both individual and sequential edits with high precision on a diverse set of images while preserving identity. Further, we propose a novel task of Attribute Style Manipulation to generate diverse styles for attributes such as eyeglass and hair. We first encode a set of synthetic images of the same identity but having different attribute styles in the latent space to estimate an attribute style manifold. Sampling a new latent from this manifold will result in a new attribute style in the generated image. We propose a novel sampling method to sample latent from the manifold, enabling us to generate a diverse set of attribute styles beyond the styles present in the training set. FLAME can generate diverse attribute styles in a disentangled manner. We illustrate the superior performance of FLAME against previous image editing methods by extensive qualitative and quantitative comparisons. FLAME generalizes well on out-of-distribution images from art domain as well as on other datasets such as cars and churches.
Rishubh Parihar, Ankit Dhiman, Tejan Karmali, Venkatesh Babu Radhakrishnan
ACM Multimedia2
2018 A Method to Generate Ghost-Free HDR Images in 360 Degree Cameras with Dual Fish-Eye Lens
abstract
360 degree cameras with fish-eye lenses recently observed increase in popularity due to increasing demand of 360 degree content. However, limitation of conventional lenses to capture entire dynamic range of scene restricts these cameras to produce an image with a high dynamic range (HDR). This is a major problem for ultra-wide angle lenses such as a fish-eye lens because larger field of view tends to have large brightness range. This paper discusses the challenges that occur when existing HDR methods are used as it is on fish-eye images, and proposes a novel method to generate a 360 degree HDR image which is fast and efficient. Secondly, this paper proposes a solution to mitigate the brightness artifact that occurs in the overlapped region of two fish-eye views. We prove the efficacy of the algorithm by comparing the generated HDR output with standard normal HDR and 360 degree HDR datasets created by us.
Ankit Dhiman, Jayakrishna Alapati, Sankaranarayanan Parameswaran, Eunsunb Ahn
ICME1