Puneet Goyal

dblp:51/10700 · DBLP profile ↗
← Back
23ranked-venue papers
3as first author
19since 2021 · last 2026
0000-0002-6196-9347ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MOT-STM: Maritime Object Tracking: A Spatial-Temporal and Metadata-based approach
Vinayak Nageli, Puneet Goyal, Rama Krishna Sai S. Gorthi
Image Vis. Comput.3
2026 AF-CANet: Augmentation-free, curriculum-guided attention UNet framework for few-shot document layout analysis
Hadia Showkat Kawoosa, Sahana Rangasrinivasan, Srirangaraj Setlur, Puneet Goyal, Venu Govindaraju
Pattern Recognit. Lett.4
2025 Multimodal Fusion Learning with Dual Attention for Medical Imaging
abstract
Multimodal fusion learning has shown significant promise in classifying various diseases such as skin cancer and brain tumors. However, existing methods face three key limitations. First, they often lack generalizability to other diagnosis tasks due to their focus on a particular disease. Second, they do not fully leverage multiple health records from diverse modalities to learn robust complementary information. And finally, they typically rely on a single attention mechanism, missing the benefits of multiple attention strategies within and across various modalities. To address these issues, this paper proposes a dual robust information fusion attention mechanism (DRIFA) that leverages two attention modules - i.e., multi-branch fusion attention module and the multimodal information fusion attention module. DRIFA can be integrated with any deep neural network, forming a multimodal fusion learning framework denoted as DRIFA-Net. We show that the multi-branch fusion attention of DRIFA learns enhanced representations for each modality, such as dermoscopy, pap smear, MRI, and CT-scan, whereas multimodal information fusion attention module learns more refined multimodal shared representations - improving the network's generalization across multiple tasks and enhancing overall performance. Additionally, to estimate the uncertainty of DRIFA-Net predictions, we have employed an ensemble Monte Carlo dropout strategy. Extensive experiments on five publicly available datasets with diverse modalities demonstrate that our approach consistently outperforms state-of-the-art methods. The code is available at https://github.com/misti1203/DRIFA-Net.
Joy Dhar, Nayyar Abbas Zaidi, Maryam Haghighat, Sudipta Roy 0002, Puneet Goyal, Azadeh Alavi
WACV5
2025 Robust page object detection network for heterogeneous document images
Hadia Showkat Kawoosa, Muhammad Suhaib Kanroo, Kapil Rana, Puneet Goyal
Int. J. Document Anal. Recognit.4
2024 Burnsnet: Burn Region Segmentation Network From Color Images With Two-Way CNN
abstract
Burn injury is a serious health issue leading to several thousands of annual fatalities. The color image-based automated burns diagnostic and assessment methods hold the potential for timely diagnosis and treatment. However, the research is limited in this domain which remains a major challenge. In this work, we explore and address the complex task of burn region segmentation in color images of burn patients. We present a semantic segmentation network that has two parallel sub-networks: a spatial-stream network for extracting low-level features and a contextual-stream network for generating a larger receptive field. Our network utilizes the pre-trained ResNet101 network, global average pooling, and instance normalization for better encoding and fusion of the network outputs. This dual-stream approach optimizes the performance in situations where data scarcity poses a challenge, facilitating robust semantic segmentation despite limited training samples. We prepared a pixel-wise labeled dataset for burn region segmentation and the experimental results on this dataset show that our proposed network outperforms several state-of-the-art semantic segmentation methods. Our method achieved mIOU and Matthews’ correlation coefficient (MCC) of 74.3% and 81.7%, respectively, approximately 4.5% higher than the second-best performing method. The Extended Burn Image Segmentation (EBIS) dataset and our model are available at https://github.com/VEDAs-Lab/EBIS
Joohi Chauhan, Paul L. Rosin, Puneet Goyal
ICIP3
2024 Uncertainty-RIFA-Net: Uncertainty Aware Robust Information Fusion Attention Network for Brain Tumors Classification in MRI Images
Joy Dhar, Kapil Rana, Puneet Goyal
ICPR (27)3
2024 PiExtract: An End-to-End Data Extraction Pipeline for Pie-Charts
Muhammad Suhaib Kanroo, Hadia Showkat Kawoosa, Joy Dhar, Puneet Goyal
ICPR (3)4
2024 Dual-branch convolutional neural network for robust camera model identification
Kapil Rana, Puneet Goyal
Expert Syst. Appl.2
2023 LYLAA: A Lightweight YOLO based Legend and Axis Analysis method for CHART-Infographics
abstract
Chart Data Extraction (CDE) is a complex task in document analysis that involves extracting data from charts to facilitate accessibility for various applications, such as document mining, medical diagnosis, and accessibility for the visually impaired. CDE is challenging due to the intricate structure and specific semantics of charts, which include elements such as title, axis, legend, and plot elements. The existing solutions for CDE have not yet satisfactorily addressed these issues. In this paper, we focus on two critical subtasks in CDE, Legend Analysis and Axis Analysis, and present a lightweight YOLO-based method for detection and domain-specific heuristic algorithms (Axis Matching and Legend Matching), for matching. We evaluate the efficacy of our proposed method, LYLAA, on a real-world dataset, the ICPR2022 UB PMC dataset, and observe promising results compared to the competing teams in the ICPR2022 CHART-Infographics competition. Our findings showcase the potential of our proposed method in the CDE process.
Hadia Showkat Kawoosa, Muhammad Suhaib Kanroo, Puneet Goyal
DocEng3
2023 A novel privacy protection approach with better human imperceptibility
Kapil Rana, Aman Pandey, Anirudh Goyal, Gurinder Singh 0001, Puneet Goyal
Appl. Intell.5
2023 IPDCN2: Improvised Patch-based Deep CNN for facial retouching detection
Kamal Sharma, Gurinder Singh 0001, Puneet Goyal
Expert Syst. Appl.3
2023 SNRCN2: Steganalysis noise residuals based CNN for source social network identification of digital images
Kapil Rana, Gurinder Singh 0001, Puneet Goyal
Pattern Recognit. Lett.3
2023 Generic multispectral demosaicking using spectral correlation between spectral bands and pseudo-panchromatic image
Vishwas Rathi, Puneet Goyal
Signal Process. Image Commun.2
2022 NCERT5K-IITRPR: A Benchmark Dataset for Non-textual Component Detection in School Books
Hadia Showkat Kawoosa, Mandhatya Singh, Manoj Manikrao Joshi, Puneet Goyal
DAS4
2022 MDCADNet: Multi dilated & context aggregated dense network for non-textual components classification in digital documents
Mandhatya Singh, Puneet Goyal
Expert Syst. Appl.2
2022 Multispectral Image Demosaicking Based on Novel Spectrally Localized Average Images
abstract
In thisletter, we propose a new generic multispectral image demosaicking algorithm using adaptive spectral correlation. Our proposed algorithm defines the spectrally localized average image (SLAI) corresponding to each spectral band, which has a strong spectral correlation with the pixel values of the corresponding spectral band in the raw image captured using a single sensor. The proposed algorithm uses the SLAI to estimate the missing pixel values of the spectral bands. Experimental results reveal that our algorithm outperforms other state-of-the-art generic multispectral image demosaicking algorithms in terms of objective and subjective evaluations.
Vishwas Rathi, Puneet Goyal
IEEE Signal Process. Lett.2
2022 SDCN2: A Shallow Densely Connected CNN for Multi-Purpose Image Manipulation Detection
abstract
Digital image information can be easily tampered with to harm the integrity of someone. Thus, recognizing the truthfulness and processing history of an image is one of the essential concerns in multimedia forensics. Numerous forensic methods have been developed by researchers with the ability to detect targeted editing operations. However, creating a unified forensic approach capable of detecting multiple image manipulations remains a challenging problem. In this article, a new general-purpose forensic approach is designed based on a shallow densely connected convolutional neural network (SDCN2) that exploits local dense connections and global residual learning. The residual domain is considered in the proposed network rather than the spatial domain to analyze the image manipulation artifacts because the residual domain is less dependent on image content information. To attain this purpose, a residual convolutional layer is employed at the beginning of the proposed model to adaptively learn the image manipulation features by suppressing the image content information. Then, the obtained image residuals or prediction error features are further processed by the shallow densely connected convolutional neural network for high-level feature extraction. In addition, the hierarchical features produced by the densely connected blocks and prediction error features are fused globally for better information flow across the network. The extensive experiment results show that the proposed scheme outperforms the existing state-of-the-art general-purpose forensic schemes even under anti-forensic attacks, when tested on large-scale datasets. The proposed model offers overall detection accuracies of 98.34% and 99.22% for BOSSBase and Dresden datasets, respectively, for multiple image manipulation detection. Moreover, the proposed network is highly efficient in terms of computational complexity as compared to the existing approaches.
Gurinder Singh 0001, Puneet Goyal
ACM Trans. Multim. Comput. Commun. Appl.2
2021 A Multi-path CNN for Automated Skin Lesion Segmentation
abstract
Automatic skin lesion segmentation in dermoscopic images is an essential requirement for making the computer-aided diagnosis (CADs) system, but efficiently segmenting skin lesions by using automated methods is not easy due to the factors such as color variations, illumination variations, presence of hair, etc. Researchers have recently been exploring deep convolutional neural networks (CNN) based methods in this domain. In this paper, we present a new and effective multi-path deep CNN method for automated skin lesion segmentation. The proposed network uses an encoder network in the first path and uses the learning representation from a pretrained base model's intermediate layers in its second path, for better learning of the features at different levels. Also, we use Instance Normalization that makes the network adaptive for each image and alleviates the problem that occurs due to different intensity images. Our method does not require any pre- or post- processing of the input dermoscopic images, except resizing in the beginning. The comparative performance evaluation of the proposed method is performed by considering two benchmark datasets: ISBI-2016 and ISBI-2017 and commonly used evaluation metrics including jaccard index and dice coefficient. The results analysis demonstrates the effectiveness of our approach, with our method achieving better performance in comparison to existing state-of-the-art methods overall and also on melanoma and non-melanoma images, across both the datasets.
Joohi Chauhan, Puneet Goyal
IJCNN2
2021 GIMD-Net: An effective General-purpose Image Manipulation Detection Network, even under anti-forensic attacks
abstract
The digital image information can be easily tampered to harm the integrity of someone. Thus, recognizing the truthfulness and processing history of an image is one of the essential concerns in multimedia forensics. Numerous forensic methods have been developed by researchers with the ability to detect targeted editing operations. But, creating a unified forensic approach capable of detecting multiple image manipulations is still a challenging problem. In this paper, a new GIMD network is designed that exploits local dense connections and global residual learning for better classification by using robust residual dense blocks (RDBs). The network input and high-level hierarchical features produced by proposed residual dense blocks are fused globally for better information flow across the network. The extensive experiment results show that the proposed scheme outperforms the existing state-of-the-art general-purpose forensic schemes even under anti-forensic attacks, when tested on large scale publicly available datasets. Our model offers overall detection accuracies of 95.09% and 97.31 % for BOSSBase and Dresden datasets, respectively for multiple image manipulation detection.
Gurinder Singh 0001, Puneet Goyal
IJCNN2
2020 Deep Learning based fully automatic efficient Burn Severity Estimators for better Burn Diagnosis
abstract
Each year, burn injuries lead to several deaths and lifelong disabilities for many others. Timely provided appropriate diagnosis and treatment can reduce sufferings for many, however automated burns diagnosis techniques are still under exploration. Laser Doppler Imaging (LDI) has been found as promising for burns depth assessment, but high costs, delays and portability issues limit its usage in developing automated burns diagnosis methods. The visual images based automated approaches for burn diagnosis have been limitedly explored. This research presents a deep learning based novel approach for burn severity assessment and a new labeled dataset of burn images with varying burn severity that would be made publically available in order to facilitate and advance research for burn severity estimation. As skin characteristics vary across different body regions so will be the burn impact, so we propose customized burn severity estimators (specific to body parts) instead of having a single generic burn severity estimator for the whole human body. Extensive experiments were conducted to evaluate the performance of the proposed approach with different network settings, obtaining competitive results to state-of-the-art methods, despite each customized estimator using a smaller set of images compared to generic one. Also, the experiments suggest that the deep learning based customized estimators perform better than handcrafted features based methods for burns diagnosis.
Joohi Chauhan, Puneet Goyal
IJCNN2
2016 Print quality assessment for stochastic clustered-dot halftones using compactness measures
abstract
Most electro-photographic printers prefer clustered-dot halftone textures for rendering smooth and stable prints. Clustered-dot halftone patterns can be periodic or aperiodic. As periodic clustered-dot halftone can lead to undesirable moiré patterns, stochastic clustered-dot halftone textures are more preferred. There are available different screening methods to generate stochastic clustered-dot halftone textures but there are no standard print quality assessment measures that can be easily used for quantitatively evaluating and comparing different stochastic clustered-dot halftoning methods. We explore the use of compactness measures for this purpose, and also propose a new compactness measure that seems good metric to quantitatively compare and assess the print quality of different stochastic clustered-dot halftoning methods. Using the proposed metric, we compare three different stochastic clustered-dot halftoning methods, and our results are almost in agreement with psychophysical experiments results reported earlier.
Puneet Goyal, Jan P. Allebach
ICIP1
2013 Clustered-Dot Halftoning With Direct Binary Search
abstract
In this paper, we present a new algorithm for aperiodic clustered-dot halftoning based on direct binary search (DBS). The DBS optimization framework has been modified for designing clustered-dot texture, by using filters with different sizes in the initialization and update steps of the algorithm. Following an intuitive explanation of how the clustered-dot texture results from this modified framework, we derive a closed-form cost metric which, when minimized, equivalently generates stochastic clustered-dot texture. An analysis of the cost metric and its influence on the texture quality is presented, which is followed by a modification to the cost metric to reduce computational cost and to make it more suitable for screen design.
Puneet Goyal, Madhur Gupta, Carl Staelin, Mani Fischer, Omri Shacham, Jan P. Allebach
IEEE Trans. Image Process.1
2011 Electro-photographic model based stochastic clustered-dot halftoning with direct binary search
abstract
Most electrophotographic printers use periodic, clustered-dot screening for rendering smooth and stable prints. However, when used for color printing, this approach suffers from the problem of periodic moire resulting from interference between the periodic halftones of individual color planes. There has been proposed an approach, called CLU-DBS for stochastic, clustered-dot halftoning and screen design based on direct binary search. We propose a methodology to embed a printer model within this halftoning algorithm to account for dot-gain and dot-loss effects. Without accounting for these effects, the printed image will not have the appearance predicted by the halftoning algorithm. We incorporate a measurement-based stochastic model for dot interactions of an electro-photographic printer within the iterative CLU-DBS binary halftoning algorithm. The stochastic model developed is based on microscopic absorptance and variance measurements. The experimental results show that electrophotography-model based stochastic clustered dot halftoning improves the homogeneity and reduces the graininess of printed halftone images.
Puneet Goyal, Madhur Gupta, Carl Staelin, Mani Fischer, Omri Shacham, Tamar Kashti, Jan P. Allebach
ICIP1