VLDB 2026 Research / reviewers in the wild / expert
Manan Shah
dblp:51/9584
· DBLP profile ↗
12ranked-venue papers
2as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Reflecting Reality: Enabling Diffusion Models to Produce Faithful Mirror ReflectionsabstractWe tackle the problem of generating highly realistic and plausible mirror reflections using diffusion-based generative models. We formulate this problem as an image inpainting task, allowing for more user control over the placement of mirrors during the generation process. To enable this, we create SynMirror, a large-scale dataset of diverse synthetic scenes with objects placed in front of mirrors. SynMirror contains around 198K samples rendered from 66K unique 3D objects, along with their associated depth maps, normal maps and instance-wise segmentation masks, to capture relevant geometric properties of the scene. Using this dataset, we propose a novel depth-conditioned inpainting method called MirrorFusion, which generates high-quality geometrically consistent and photo-realistic mirror reflections given an input image and a mask depicting the mirror region. MirrorFusion outperforms state-of-the-art methods on SynMirror, as demonstrated by extensive quantitative and qualitative analysis. To the best of our knowledge, we are the first to successfully tackle the challenging problem of generating controlled and faithful mirror reflections of an object in a scene using diffusion based models. Syn-Mirror and MirrorFusion open up new avenues for image editing and augmented reality applications for practitioners and researchers alike. The project page is available at: https://val.cds.iisc.ac.in/reflecting-reality.github.io/. Ankit Dhiman, Manan Shah, Rishubh Parihar, Yash Bhalgat, Lokesh R. Boregowda, Venkatesh Babu Radhakrishnan |
3DV | 2 |
| 2025 | MirrorVerse: Pushing Diffusion Models to Realistically Reflect the WorldabstractDiffusion models have become central to various image editing tasks, yet they often fail to fully adhere to physical laws, particularly with effects like shadows, reflections, and occlusions. In this work, we address the challenge of generating photorealistic mirror reflections using diffusion-based generative models. Despite extensive training data, existing diffusion models frequently overlook the nuanced details crucial to authentic mirror reflections. Recent approaches have attempted to resolve this by creating synthetic datasets and framing reflection generation as an in-painting task; however, they struggle to generalize across different object orientations and positions relative to the mirror. Our method overcomes these limitations by introducing key augmentations into the synthetic data pipeline: (1) random object positioning, (2) randomized rotations, and (3) grounding of objects, significantly enhancing generalization across poses and placements. To further address spatial relationships and occlusions in scenes with multiple objects, we implement a strategy to pair objects during dataset generation, resulting in a dataset robust enough to handle these complex scenarios. Achieving generalization to real-world scenes remains a challenge, so we introduce a three-stage training curriculum to develop the MirrorFusion 2.0 model to improve real-world performance. We provide extensive qualitative and quantitative evaluations to support our approach. The project page is available at: https://mirror-verse.github.io/. Ankit Dhiman, Manan Shah, Venkatesh Babu Radhakrishnan |
CVPR | 2 |
| 2025 | ContextGNN: Beyond Two-Tower Recommendation SystemsabstractRecommendation systems predominantly utilize two-tower architectures, which evaluate user-item rankings through the inner product of their respective embeddings. However, one key limitation of two-tower models is that they learn a pair-agnostic representation of users and items. In contrast, pair-wise representations either scale poorly due to their quadratic complexity or are too restrictive on the candidate pairs to rank. To address these issues, we introduce Context-based Graph Neural Networks (ContextGNNs), a novel deep learning architecture for link prediction in recommendation systems. The method employs a pair-wise representation technique for familiar items situated within a user's local subgraph, while leveraging two-tower representations to facilitate the recommendation of exploratory items. A final network then predicts how to fuse both pair-wise and two-tower recommendations into a single ranking of items. We demonstrate that ContextGNN is able to adapt to different data characteristics and outperforms existing methods, both traditional and GNN-based, on a diverse set of practical recommendation tasks, improving performance by 20\% on average. Yiwen Yuan, Zecheng Zhang, Akihiro Nitta, Weihua Hu, Manan Shah, Blaz Stojanovic, Shenyang Huang, Jan Eric Lenssen, Jure Leskovec, Matthias Fey |
ICLR | 6 |
| 2025 | Crop protection and disease detection using artificial intelligence and computer vision: a comprehensive review
Kanish Shah, Rajat Sushra, Manan Shah, Haard Shah, Megh Raval, Mitul Prajapati |
Multim. Tools Appl. | 3 |
| 2025 | Advanced driver assistance system (ADAS) and machine learning (ML): The dynamic duo revolutionizing the automotive industryabstractThe advanced driver assistance system (ADAS) primarily serves to assist drivers in monitoring the speed of the car and helps them make the right decision, which leads to fewer fatal accidents and ensures higher safety. In the artificial Intelligence domain, machine learning (ML) was developed to make inferences with a degree of accuracy similar to that of humans; however, enormous amounts of data are required. Machine learning enhances the accuracy of the decisions taken by ADAS, by evaluating all the data received from various vehicle sensors. This study summarizes all the critical algorithms used in ADAS technologies and presents the evolution of ADAS technology. Initially, ADAS technology is introduced, along with its evolution, to understand the objectives of developing this technology. Subsequently, the critical algorithms used in ADAS technology, which include face detection, head-pose estimation, gaze estimation, and link detection are discussed. A further discussion follows on the impact of ML on each algorithm in different environments, leading to increased accuracy at the expense of additional computing, to increase efficiency. The aim of this study was to evaluate all the methods with or without ML for each algorithm. Karan Shah 0005, Kushagra Darji, Adit Shah, Manan Shah |
Virtual Real. Intell. Hardw. | 5 |
| 2024 | All phase discrete cosine biorthogonal transform versus discrete cosine transform in digital watermarking
Dev Tailor, Kevin Panchal, Samir Patel, Manan Shah |
Multim. Tools Appl. | 6 |
| 2023 | A comprehensive study on lane detecting autonomous car using computer vision
Henil Gajjar, Stavan Sanyal, Manan Shah |
Expert Syst. Appl. | 3 |
| 2022 | A stock market trading framework based on deep learning architectures
Atharva Shah, Maharshi Gor, Meet Sagar, Manan Shah |
Multim. Tools Appl. | 4 |
| 2021 | A comprehensive analysis on movie recommendation system employing collaborative filtering
Urvish Thakker, Ruhi Patel, Manan Shah |
Multim. Tools Appl. | 3 |
| 2020 | Machine Learning to Identify Peripherally Inserted Central Catheter (PICC) Tip Position from Radiology Reports
Manan Shah, Derek Shu, V. B. Surya Prasath, Yizhao Ni, Andrew Schapiro, Kevin R. Dufendach |
AMIA | 1 |
| 2019 | Inferring Context from Pixels for Multimodal Image ClassificationabstractImage classification models take image pixels as input and predict labels in a predefined taxonomy. While contextual information (e.g. text surrounding an image) can provide valuable orthogonal signals to improve classification, the typical setting in literature assumes the unavailability of text and thus focuses on models that rely purely on pixels. In this work, we also focus on the setting where only pixels are available in the input. However, we demonstrate that if we predict textual information from pixels, we can subsequently use the predicted text to train models that improve overall performance. We propose a framework that consists of two main components: (1) a phrase generator that maps image pixels to a contextual phrase, and (2) a multimodal model that uses textual features from the phrase generator and visual features from the image pixels to produce labels in the output taxonomy. The phrase generator is trained using web-based query-image pairs to incorporate contextual information associated with each image and has a large output space. We evaluate our framework on diverse benchmark datasets (specifically, the WebVision dataset for evaluating multi-class classification and OpenImages dataset for evaluating multi-label classification), demonstrating performance improvements over approaches based exclusively on pixels and showcasing benefits in prediction interpretability. We additionally present results to demonstrate that our framework provides improvements in few-shot learning of minimally labeled concepts. We further demonstrate the unique benefits of the multimodal nature of our framework by utilizing intermediate image/text co-embeddings to perform baseline zero-shot learning on the ImageNet dataset. Manan Shah, Krishnamurthy Viswanathan, Chun-Ta Lu, Ariel Fuxman, Zhen Li 0028, Aleksei Timofeev, Chao Jia 0005, Chen Sun 0002 |
CIKM | 1 |
| 2015 | The design and control of the multi-modal locomotion origami robot, TribotabstractOrigami robots (Robogamis) use architecture to strategically activate different sets and sequence of actuators to achieve large variety of reconfigurable forms. Tribot is a unique mobile origami robot that can simultaneously choose between two modes of locomotion: jumping and crawling. When assembled, Tribot measures 64 × 34 × 20 mm3, weighs 4 g, crawls at 17% of its body length per gait cycle and jumps seven times its height repeatedly without needing to be reset. To optimize the practicality of the nominally 2D design, we made two different approaches to build the prototypes. For one of them, we used the “traditional”, monolithic, layer-by-layer robogami fabrication method and the second, we printed out most parts using a multi-material 3D printer. By showing the performance of two prototypes side-by-side, we show that with the 3D printer, we can minimize the number of functional layers and reduce the fabrication time. The embedded sensors allow Tribot's crawling gait pattern and jumping height to be modulated with a closed loop control. We compare the expected gait step size and displacement to that of the presented prototype while describing the design and control parameters to achieve the experimental results. We also illustrated the preliminary graphical design tool platform developed to optimize the next design iteration of Tribot. Zhenishbek Zhakypov, Mohsen Falahi, Manan Shah, Jamie Kyujin Paik |
IROS | 3 |