VLDB 2026 Research / reviewers in the wild / expert
Chiranjoy Chattopadhyay
dblp:125/2929
· DBLP profile ↗
32ranked-venue papers
3as first author
21since 2021 · last 2026
0000-0002-3431-0483ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Approach Towards Developing Relationally Intelligent Multimodal Framework for Stock Movement Prediction (Student Abstract)abstractThe dependency of stock prices on a multitude of factors makes the task of prediction exceedingly challenging. Given the volatile nature of stock data, it is imperative to integrate multiple sources of information to accurately encompass the various factors that influence market trends. To capture these complex dynamics, several multimodal methodologies have been proposed, integrating market data, technical indicators, and textual information. However, it is claimed that these coarse-grained information sources do not offer a holistic view of the market. Furthermore, these sources are stock-specific and do not elucidate the interconnections between various stocks. To address this deficiency, we propose a multimodal approach that incorporates this relational aspect alongside fine-grained information sources. The applicability of our framework is underscored by empirical results, which demonstrate the superiority of our approach. Manali Patel, Krupa Jariwala, Chiranjoy Chattopadhyay |
AAAI | 3 |
| 2026 | Multi-modal Heterogeneous Music Notation Recognition from Documents: A Framework for Real-Time Symbolic-to-Immersive Mapping
Ankit Sinha, Atanu Saha, Chiranjoy Chattopadhyay, Rahul Kumar Ray |
ICDAR (2) | 3 |
| 2026 | A Comprehensive Survey of Floor Plan Image Analysis and Related Applications
Rasika Khade, Krupa Jariwala, Chiranjoy Chattopadhyay |
Int. J. Document Anal. Recognit. | 3 |
| 2026 | Advancing abstract reasoning for RPMs with a path aggregation network and deep predictive reasoning
Amresh Kumar Singh, Sandeep Khanna, Chiranjoy Chattopadhyay |
Mach. Vis. Appl. | 3 |
| 2026 | Watermarking Protocol Inspired Kidney Stone Segmentation in IoMTabstractThe rapid explosion of medical data, exarcebated by the demands of smart healthcare, poses significant challenges for authentication and integrity verification. Moreover, the surge in cybercrime targeting healthcare data jeopardizes patient privacy, compromising both trust and diagnostic reliability. To address these concerns, we propose a robust healthcare system that integrates a kidney stone segmentation framework with a watermarking protocol tailored for Internet of Medical Things (IoMT) applications. Drawing upon patient information and biometrics, chaotic keys are generated for obfuscation and randomization, along with the watermark for integrity verification and authentication. The watermark is imperceptibly embedded into the obfuscated medical image using Singular Value Decomposition (SVD) and adaptive quantization, followed by randomization. Upon reception, successful watermark extraction and verification ensure secure access to unaltered medical data, enabling precise segmentation. To facilitate this, a ResNeXt-50 inspired encoder and attention-guided decoder are introduced within the U-Net architecture to enhance comprehensive feature learning. The effectiveness and practicality of the proposed system have been evaluated through comprehensive experiments on kidney CT scans. Comparative analysis with state-of-the-art techniques highlights its superior performance. Parkala Vishnu Bharadwaj Bayari, Nishtha Tomar, Gaurav Bhatnagar, Chiranjoy Chattopadhyay |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | From Notes to Keys: A VR Learning Environment for Sheet Music Interpretation
Sandeep Khanna, Atanu Saha, Rahul Kumar Ray, Rakesh Patibanda, Chiranjoy Chattopadhyay |
ICDAR (3) | 5 |
| 2025 | A multi relational graph-based approach for stock movement prediction and portfolio optimization
Manali Patel, Shreya Goyal, Krupa Jariwala, Chiranjoy Chattopadhyay |
Neurocomputing | 4 |
| 2025 | A Novel Chaotic Map and Its Application to Secure Transmission of Multimodal ImagesabstractThe advent of digital technology, augmented by connected devices, has catalyzed a dramatic increase in multimedia content consumption, facilitating on-the-go access and communication. However, this surge also heightens the risks of unauthorized access, privacy breaches, and cyberattacks. Consequently, ensuring the secure and efficient transmission and storage of multimedia content is of paramount importance. This article presents a robust encryption scheme for secure image transmission, utilizing a novel one-dimensional chaotic map characterized by random and complex dynamics, validated through NIST test and meticulous evaluation. Key matrices are derived from the chaotic map, with the SHA-256 hash of random, nonoverlapping blocks of the input image influencing the initial conditions, thereby ensuring resistance to differential cryptanalysis. The encryption process encompasses a dual shuffling mechanism: an adaptive shuffling guided by the chaotic key, followed by orbital shuffling, which rearranges pixel positions by segmenting the image into distinct orbital patterns. This is complemented by a feedback diffusion technique that ensures each pixel’s encryption is influenced by neighboring values and the keys employed. Extensive evaluation with multimodal images demonstrates the scheme’s versatility, with significant resilience against various cryptographic attacks, as evidenced by thorough assessments. Comparative analysis further highlights the superiority of the proposed scheme over state-of-the-art approaches. These attributes position the proposed scheme as a highly effective solution for contemporary digital security challenges. Parkala Vishnu Bharadwaj Bayari, Yashmita Sangwan, Gaurav Bhatnagar, Chiranjoy Chattopadhyay |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2025 | An Approach Toward Stock Market Prediction and Portfolio Optimization in Indian Financial SectorsabstractIn this article, we aim at predicting future stock price movements and recommending a profitable portfolio for the NIFTY-50 stocks. Stock market prediction is a challenging task due to multiple influencing factors, its nonlinear and volatile nature, and complex interdependencies. Recent approaches have neglected the interconnections between stocks and relied on predefined static relationships. The collection of relational data is difficult to access due to confidentiality and privacy agreements for emerging economies. Moreover, these predefined relationships lack the ability to explain the latent interactions between stocks. This work proposes a data-driven end-to-end framework, dynamic relation aware relational temporal network (DR2TNet), that learns the hidden intra- and intersector associations between stock pairs and temporal patterns. A financial knowledge graph is built from historical data and is updated dynamically during the training process to reflect the interactions between the stocks according to the current market situation. We have proposed a new loss function that considers prediction loss and directional movement loss to train a model. The applicability of prediction results obtained by DR2TNet is demonstrated in the portfolio optimization problem. The results show a higher return compared to other existing baseline models. Manali Patel, Krupa Jariwala, Chiranjoy Chattopadhyay |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | Synergizing Vision and Language in Remote Sensing: A Multimodal Approach for Enhanced Disaster Classification in Emergency Response SystemsabstractAs remote sensing capabilities continue to advance, there is a growing interest in leveraging computer vision and natural language processing for enhanced interpretation of remote sensing scenes. This paper explores the integration of textual information with images to augment traditional disaster classification methods. Our approach utilizes a predefined vision-language model to generate descriptive captions for images, fostering a more nuanced understanding of the remote sensing data. Next, we seamlessly integrate the generated textual information with image data through multimodal training, employing a multimodal deep learning method for disaster classification. The system categorizes input data into predefined disaster categories, presenting a comprehensive and accurate approach to emergency response system development. Experimental evaluations conducted on the AIDER dataset (Aerial Image Database for Emergency Response applications) showcase the efficacy of our approach, demonstrating improved accuracy compare to unimodal approach and reliability in disaster classification. This research contributes to the advancement of intelligent emergency response systems by harnessing the synergy between vision and language in the context of remote sensing. Nandini Saini, Suman Kundu, Chiranjoy Chattopadhyay, Debasis Das 0001 |
IGARSS | 4 |
| 2024 | EVDNET: Towards Explainable Multi Scale, Anchor Free Vehicle Detection Network in High Resolution Aerial ImageryabstractThe rapid advancement in deep learning-based object detection methods has made them a prevalent choice for real-time applications. Families of object detectors, including one-stage detectors, two-stage detectors, and region-based CNN networks, offer superior performance in accurately detecting objects. Despite their high accuracy, the complex design and black-box functionality of these models are not directly transferable in aerial imagery. Also, raise questions among users regarding the transparency of the algorithm in locating objects. Consequently, to demystify the decision process of these models, there is a need for Explainable AI (XAI) tools. XAI enables an understanding of the significance of each pixel in an image, shedding light on the contributions that lead to the model’s final output. In this context, this work will present an efficient, explainable, multi-scale vehicle detection network from high resolution aerial imagery, named as EVDNet. The EVDNet model has trained with two publicly available aerial image benchmark dataset DOTA and VEDAI. To enhance interpretability, we leverage XAI method using GradCam. The experimental results not only showcase the effectiveness and performance of the EVDNet model but also provide valuable insights into the object detection process. This research contributes to bridging the gap between complex object detection models and user understanding, offering a more transparent and interpretable approach to high-resolution aerial imagery analysis. Nandini Saini, Chiranjoy Chattopadhyay, Debasis Das 0001, Suman Kundu |
IGARSS | 3 |
| 2023 | DG-YOLOT: A Lightweight Density Guided YOLO-Transformer for Remote Sensing Object DetectionabstractDeep learning-based object detection methods in natural image datasets have demonstrated remarkable accuracy and lower error rates than those of humans. As a result, they have gained significant attention in the field of remote sensing imagery. However, direct transferablity of these methods in remote sensing images face challenges such as scale variations, complex object distributions, and arbitrary orientations. In order to address these challenges, we propose the transformer-based object detector named as DG-YOLOT where we use a guided self-attention mechanism with YOLOv5 to enhance the potentiality of training with minimum computation. Instead of employing uniform size patches like the conventional vision transformer, we leverage density map patches which facilitates the extraction of diverse contextual information related to objects within the image, enhancing the differentiation capability of our model. Through extensive experiments conducted on the DOTAv2.0 dataset, our proposed model has demonstrated superior performance with 57.91% mean Average Precision (mAP) compared to other state-of-the-art object detectors. Nandini Saini, Chiranjoy Chattopadhyay, Debasis Das 0001 |
IGARSS | 2 |
| 2023 | On enhancing prediction abilities of vision-based metallic surface defect classification through adversarial training
Vikanksh Nath, Chiranjoy Chattopadhyay, K. A. Desai |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | SOLARNet: A single stage regression based framework for efficient and robust object recognition in aerial images
Nandini Saini, Chiranjoy Chattopadhyay, Debasis Das 0001 |
Pattern Recognit. Lett. | 2 |
| 2022 | An Ontological Approach towards Automatic Creation of Infographics from Formal Text (Student Abstract)abstractInfographics deal with representing data or information visually in a perceptually compelling manner. Recently, infographics have gained widespread popularity, giving rise to automated infographics synthesis from texts. Our research follows an ontological approach to automatically extract the necessary indicators from an input sentence and synthesize an infographic corresponding to it. This work includes (1) the creation of a dataset, (2) an end-to-end domain-agnostic framework, and (3) demonstrating the application of the proposed framework. The results demonstrate our framework's ability to extract the necessary textual cues from real-world textual descriptions (from various domains) and synthesize meaningful infographics. Devin Garg, Tanuj Agarwal, Chiranjoy Chattopadhyay |
AAAI | 3 |
| 2022 | GRIHA: synthesizing 2-dimensional building layouts from images captured using a smart phone
Shreya Goyal, Naimul Mefraz Khan, Chiranjoy Chattopadhyay, Gaurav Bhatnagar |
Multim. Tools Appl. | 3 |
| 2022 | RISC-Net : rotation invariant siamese convolution network for floor plan image retrieval
Atharva Kalsekar, Rasika Khade, Krupa Jariwala, Chiranjoy Chattopadhyay |
Multim. Tools Appl. | 4 |
| 2021 | C2VNet: A Deep Learning Framework Towards Comic Strip to Audio-Visual Scene Synthesis
Vaibhavi Gupta, Vinay Detani, Vivek Khokar, Chiranjoy Chattopadhyay |
ICDAR (2) | 4 |
| 2021 | S2D2Net: An Improved Approach For Robust Steel Surface Defects Diagnosis With Small Sample LearningabstractSurface defect recognition of products is a necessary process to guarantee the quality of industrial production. This paper proposes a hybrid model, S2D2Net (Steel Surface Defect Diagnosis Network), for an efficient and robust inspection of the steel surface during the manufacturing process. The S2D2Net uses a pretrained ImageNet model as a feature extractor and learns a Capsule Network over the extracted features. The experimental results on a publicly available steel surface defect dataset (NEU) show that S2D2Net achieved 99.17% accuracy with minimal training data and improved by 9.59% over its closest competitor based on GAN. S2D2Net proved its robustness by achieving 94.7% accuracy on a diversity enhanced dataset, ENEU, and improved by 3.6% over its closest competitor. It has better, robust recognition performance compared to other state-of-the-art DNN-based detectors. Vikanksh Nath, Chiranjoy Chattopadhyay |
ICIP | 2 |
| 2021 | Knowledge-driven description synthesis for floor plan interpretation
Shreya Goyal, Chiranjoy Chattopadhyay, Gaurav Bhatnagar |
Int. J. Document Anal. Recognit. | 2 |
| 2021 | A rotation and scale invariant approach for multi-oriented floor plan image retrieval
Rasika Khade, Krupa Jariwala, Chiranjoy Chattopadhyay, Umapada Pal 0001 |
Pattern Recognit. Lett. | 3 |
| 2019 | Radial Loss for Learning Fine-grained Video Similarity MetricabstractIn this paper, we propose the Radial Loss which utilizes category and sub-category labels to learn an order-preserving fine-grained video similarity metric. We propose an end-to-end quadlet-based Convolutional Neural Network (CNN) combined with Long Short-term Memory (LSTM) Unit to model video similarities by learning the pairwise distance relationships between samples in a quadlet generated using the category and sub-category labels. We showcase two novel applications of learning a video similarity metric - (i) fine-grained video retrieval, (ii) fine-grained event detection, along with simultaneous shot boundary detection, and correspondingly show promising results against those of the baselines on two new fine-grained video datasets. Abhinav Jain 0001, Prerna Agarwal, Shashank Mujumdar, Nitin Gupta 0005, Sameep Mehta, Chiranjoy Chattopadhyay |
ICASSP | 6 |
| 2019 | BRIDGE: Building Plan Repository for Image Description Generation, and EvaluationabstractIn this paper, a large scale public dataset containing floor plan images and their annotations is presented. BRIDGE (Building plan Repository for Image Description Generation, and Evaluation) dataset contains more than 13000 images of the floor plan and annotations collected from various websites, as well as publicly available floor plan images in the research domain. The images in BRIDGE also has annotations for symbols, region graphs, and paragraph descriptions. The BRIDGE dataset will be useful for symbol spotting, caption and description generation, scene graph synthesis, retrieval and many other tasks involving building plan parsing. In this paper, we also present an extensive experimental study for tasks like furniture localization in a floor plan, caption and description generation, on the proposed dataset showing the utility of BRIDGE. Shreya Goyal, Vishesh Mistry, Chiranjoy Chattopadhyay, Gaurav Bhatnagar |
ICDAR | 3 |
| 2019 | SUGAMAN: describing floor plans for visually impaired by annotation learning and proximity-based grammarabstractIn this study, the authors propose a framework SUGAMAN (Supervised and Unified framework using Grammar and Annotation Model for Access and Navigation). SUGAMAN is a Hindi word meaning ‘easy passage from one place to another’. SUGAMAN synthesises textual description from a given floor plan image, usable by visually impaired to navigate by understanding the arrangement of rooms and furniture. It is the first framework for describing a floor plan and giving direction for obstacle‐free movement within a building. The model learns five classes of room categories from 1355 room image samples under a supervised learning paradigm. These learned annotations are fed into a description synthesis framework to yield a holistic description of a floor plan image. Authors demonstrate the performance of various supervised classifiers on room learning and provided a comparative analysis of system generated and human‐written descriptions. The contribution of this study includes a novel framework for description generation from document images with graphics while proposing a new feature representing the floor plans, text annotations for a publicly available data set, and an algorithm for door to door obstacle avoidance navigation. This work can be applied to areas like understanding floor plans and design of historical monuments, and retrieval. Shreya Goyal, Satya Bhavsar, Shreya Patel, Chiranjoy Chattopadhyay, Gaurav Bhatnagar |
IET Image Process. | 4 |
| 2019 | A novel feature transform framework using deep neural network for multimodal floor plan retrieval
Nitin Gupta 0005, Chiranjoy Chattopadhyay, Sameep Mehta |
Int. J. Document Anal. Recognit. | 3 |
| 2018 | REXplore: A Sketch Based Interactive Explorer for Real Estates Using Building Floor Plan ImagesabstractThe increasing trend of using online platforms for real estate rent/sale makes automatic retrieval of similar floor plans a key requirement to help architects and buyers alike. Although sketch based image retrieval has been explored in the multimedia community, the problem of hand-drawn floor plan retrieval has been less researched in the past. In this paper, we propose REXplore (Real Estate eXplore), a novel framework that uses sketch based query mode to retrieve corresponding similar floor plan images from a repository using Cyclic Generative Adversarial Networks (Cyclic GAN) for mapping between sketch and image domain. The key contributions of our proposed approach are : (1) a novel sketch based floor plan retrieval framework using an intuitive and convenient sketch query mode; (2) A conjunction of Cyclic GANs and Convolution Neural Networks (CNNs) for the task of hand-drawn floor plan image retrieval. Extensive experimentation and comparison with baseline results authenticates our claim. Nitin Gupta 0005, Chiranjoy Chattopadhyay, Sameep Mehta |
ISM | 3 |
| 2018 | High-level feature aggregation for fine-grained architectural floor plan retrievalabstractDue to the massive growth of real estate industry, there is an increase in the number of online platforms designed for finding homes/furnished properties. Instead of descriptive words, query by example is always a preferred method for retrieval. Floor plans are the basic 2D representation giving an idea about the building structure at a particular level. The authors propose a framework for the retrieval of similar architectural floor plans under the query by example paradigm. They propose a novel algorithm to extract high‐level semantic features from an architectural floor plan. Fine‐grained retrieval using weighted sum of the features is proposed, where a feature can be given more preference over others, during retrieval. Experiments were performed on publicly available dataset containing 510 floor plans and compared with existing state‐of‐the‐art techniques. Their proposed method outperforms others both in qualitative and quantitative terms. Chiranjoy Chattopadhyay |
IET Comput. Vis. | 2 |
| 2017 | DANIEL: A Deep Architecture for Automatic Analysis and Retrieval of Building Floor PlansabstractAutomatically finding out existing building layouts from a repository is always helpful for an architect to ensure reuse of design and timely completion of projects. In this paper, we propose Deep Architecture for fiNdIng alikE Layouts (DANIEL). Using DANIEL, an architect can search from the existing projects repository of layouts (floor plan), and give accurate recommendation to the buyers. DANIEL is also capable of recommending the property buyers, having a floor plan image, the corresponding rank ordered list of alike layouts. DANIEL is based on the deep learning paradigm to extract both low and high level semantic features from a layout image. The key contributions in the proposed approach are: (i) novel deep learning framework to retrieve similar floor plan layouts from repository; (ii) analysing the effect of individual deep convolutional neural network layers for floor plan retrieval task; and (iii) creation of a new complex dataset ROBIN (Repository Of BuildIng plaNs), having three broad dataset categories with 510 real world floor plans.We have evaluated DANIEL by performing extensive experiments on ROBIN and compared our results with eight different state-of-the-art methods to demonstrate DANIEL's effectiveness on challenging scenarios. Nitin Gupta 0005, Chiranjoy Chattopadhyay, Sameep Mehta |
ICDAR | 3 |
| 2016 | A unified framework for semantic matching of architectural floorplansabstractAn automatic lookup tool, which matches and retrieves similar floorplans from a large repository of digitized architectural floorplans can prove to be of immense help for the architects while designing new projects. In this paper, we have proposed a framework for the matching and retrieval of similar architectural floorplans under the query by example paradigm. We propose a room layout segmentation and adjacent room detection algorithm to represent layouts as an undirected graph. We have also proposed a novel graph spectral embedding feature to uniquely represent the layout of the architectural floorplan. This helps in effective and efficient matching of the room layouts. Room semantics in terms of both the room structures and room decor is used to retrieve similar floorplans from the repository. To match the semantic similarity between a pair of floorplans, we have proposed a two stage matching technique. We have validated the effectiveness of our proposed framework by performing experiments on publicly available floorplan dataset and achieved high retrieval accuracy. Chiranjoy Chattopadhyay, Gaurav Harit |
ICPR | 2 |
| 2016 | Supervised framework for automatic recognition and retrieval of interaction: a framework for classification and retrieving videos with similar human interactionsabstractThis study presents supervised framework for automatic recognition and retrieval of interactions (SAFARRIs), a supervised learning framework to recognise interactions such as pushing, punching, and hugging, between a pair of human performers in a video shot. The primary contribution of the study is to extend the vectors of locally aggregated descriptors (VLADs) as a compact and discriminative video encoding representation, to solve the complex class partitioning problem of recognising human interaction. An initial codebook is generated from the training set of video shots, by extracting feature descriptors around the spatiotemporal interest points computed across frames. A bag of action words is generated by encoding the first‐order statistics of the visual words using VLAD. Support vector machine classifiers (1 against all) are trained using these codebooks. The authors have verified SAFARRI's accuracy for classification and retrieval (query by example). SAFARRI is free from tracking or recognition of body parts and capable of identifying the region of interaction in video shots. It gives superior retrieval and classification performances over recently proposed methods, on two publicly available human interaction datasets. Chiranjoy Chattopadhyay, Sukhendu Das |
IET Comput. Vis. | 1 |
| 2012 | Enhancing the MST-CSS Representation Using Robust Geometric Features, for Efficient Content Based Video Retrieval (CBVR)abstractMulti-Spectro-Temporal Curvature Scale Space (MST-CSS) had been proposed as a video content descriptor in an earlier work, where the peak and saddle points were used for feature points. But these are inadequate to capture the salient features of the MST-CSS surface, producing poor retrieval results. To overcome these, we propose EMST-CSS (Enhanced MST-CSS) as a better feature representation with an improved matching method for CBVR (Content Based Video Retrieval). Comparative study with the existing MST-CSS representation and two state-of-the-art methods for CBVR shows enhanced performance on one synthetic and two real-world datasets. Chiranjoy Chattopadhyay, Sukhendu Das |
ISM | 1 |
| 2012 | A Motion-Sketch Based Video Retrieval Using MST-CSS RepresentationabstractIn this work, we propose a framework for a robust Content Based Video Retrieval (CBVR) system with free hand query sketches, using the Multi-Spectro Temporal-Curvature Scale Space (MST-CSS) representation. Our designed interface allows sketches to be drawn to depict the shape of the object in motion and its trajectory. We obtain the MST-CSS feature representation using these cues and match with a set of MST-CSS features generated offline from the video clips in the database (gallery). Results are displayed in rank ordered similarity. Experimentation with benchmark datasets shows promising results. Chiranjoy Chattopadhyay, Sukhendu Das |
ISM | 1 |