Nikhil Kumar Tomar

dblp:282/4350 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 MDNet: Multi-Decoder Network for Abdominal CT Organs Segmentation
abstract
Accurate segmentation of organs from abdominal CT scans is essential for clinical applications such as diagnosis, treatment planning, and patient monitoring. To handle challenges of heterogeneity in organ shapes, sizes, and complex anatomical relationships, we propose a Multi decoder network (MDNet), an encoder-decoder network that uses the pre-trained MiT-B2 as the encoder and multiple different decoder networks. Each decoder network is connected to a different part of the encoder via a multi-scale feature enhancement dilated block. With each decoder, we increase the depth of the network iteratively and refine segmentation masks, enriching feature maps by integrating previous decoders’ feature maps. To refine the feature map further, we also utilize the predicted masks from the previous decoder to the current decoder to provide spatial attention across foreground and background regions. MDNet effectively refines the segmentation mask with a high dice similarity coefficient (DSC) of 0.9013 and 0.9169 on the Liver Tumor segmentation (LiTS) and MSD Spleen datasets. Additionally, it reduces Hausdorff distance (HD) to 3.79 for the LiTS dataset and 2.26 for the spleen segmentation dataset, underscoring the precision of MDNet in capturing the complex contours. Moreover, MDNet is more interpretable and robust compared to the other baseline models. The code for our architecture is available at https://github.com/DebeshJha/MDNet.
Debesh Jha, Nikhil Kumar Tomar, Koushik Biswas, Gorkem Durak, Matthew Antalek, Zheyuan Zhang 0001, Bin Wang 0068, Md Mostafijur Rahman, Hongyi Pan, Alpay Medetalibeyoglu, Vandan Gorade, Yury Velichko, Daniela P. Ladner, Amir Borhani, Ulas Bagci
ICASSP2
2025 Transformer-Enhanced Iterative Feedback Mechanism For Polyp Segmentation
abstract
Colorectal cancer (CRC) is the third most common cause of cancer diagnosed in the United States. Notably, CRC is the leading cause of cancer in younger men less than 50 years old. Colonoscopy is considered the gold standard for the early diagnosis of CRC. Skills vary significantly among endoscopists, and a high miss rate is reported. Automated polyp segmentation can reduce the missed rates, and timely treatment is possible in the early stage. To address this challenge, we introduce Feedback Attention Network-v2 (FANetv2), an advanced encoder-decoder network designed to accurately segment polyps from colonoscopy images. Leveraging an initial input mask generated by Otsu thresholding, FANetv2 iteratively refines its binary segmentation masks through a novel feedback attention mechanism informed by the mask predictions of previous epochs. Additionally, it employs a text-guided approach that integrates essential information about the number (one or many) and size (small, medium, large) of polyps to further enhance its feature representation capabilities. This dual-task approach facilitates accurate polyp segmentation and aids in the auxiliary classification of polyp attributes, significantly boosting the model’s performance. Our comprehensive evaluations on the publicly available BKAI-IGH and CVC-ClinicDB datasets demonstrate the superior performance of FANetv2, evidenced by high dice similarity coefficients (DSC) of 0.9186 and 0.9481, along with low Hausdorff distances of 2.83 and 3.19, respectively. The source code for FANetv2 will be made available at https://github.com/nikhilroxtomar/FANetv2.
Nikhil Kumar Tomar, Debesh Jha, Koushik Biswas, Ulas Bagci
ICASSP1
2025 Optimizing Neural Network Effectiveness via Non-monotonicity Refinement
abstract
Activation functions play a crucial role in artificial neural networks by introducing non-linearities that enable networks to learn complex patterns in data. An appropriate choice of an activation function plays a crucial role in the training dynamics of a neural network, which can boost network performance significantly. Rectified Linear Unit (ReLU) and its variants, like leaky ReLU and parametric ReLU, have emerged as the most popular activations due to their ability to enable faster training and generalization in deep neural networks despite having some significant issues like vanishing gradient problems. In this paper, we have proposed smooth functions, which we call the AMSU family, which are smooth approximations of the maximum function. We derive three activations from the AMSU family, namely AMSU-1, AMSU-2, & AMSU-3, and show their effectiveness in different deep learning problems. By simply replacing the ReLU function, Top-1 accuracy improves by 5.88%, 5.96%, and 5.32% on the CIFAR100 dataset on the ShuffleNet V2 model. Also, replacing ReLU with AMSU-1, AMSU-2, and AMSU-3, Top-1 accuracy improves by 8.50%, 8.29%, and 7.70% on the CIFAR100 dataset on the ShuffleNet V2 model with FGSM attack. Also, Replacing ReLU with AMSU-1, AMSU-2, and AMSU-3 on ImageNet-1K data, we got 3%-5% improvement on ShuffleNet and MobileNet models. The source code is publicly available at https://github.com/koushik313/AMSU.
Koushik Biswas, Amit Reza, Meghana Karri, Debesh Jha, Hongyi Pan, Nikhil Kumar Tomar, Aliza Subedi, Smriti Regmi, Ulas Bagci
WACV6
2025 Validating polyp and instrument segmentation methods in colonoscopy through Medico 2020 and MedAI 2021 Challenges
abstract
Automatic analysis of colonoscopy images has been an active field of research motivated by the importance of early detection of precancerous polyps. However, detecting polyps during the live examination can be challenging due to various factors such as variation of skills and experience among the endoscopists, lack of attentiveness, and fatigue leading to a high polyp miss-rate. Therefore, there is a need for an automated system that can flag missed polyps during the examination and improve patient care. Deep learning has emerged as a promising solution to this challenge as it can assist endoscopists in detecting and classifying overlooked polyps and abnormalities in real time, improving the accuracy of diagnosis and enhancing treatment. In addition to the algorithm’s accuracy, transparency and interpretability are crucial to explaining the whys and hows of the algorithm’s prediction. Further, conclusions based on incorrect decisions may be fatal, especially in medicine. Despite these pitfalls, most algorithms are developed in private data, closed source, or proprietary software, and methods lack reproducibility. Therefore, to promote the development of efficient and transparent methods, we have organized the “Medico automatic polyp segmentation (Medico 2020)” and “MedAI: Transparency in Medical Image Segmentation (MedAI 2021)” competitions. The Medico 2020 challenge received submissions from 17 teams, while the MedAI 2021 challenge also gathered submissions from another 17 distinct teams in the following year. We present a comprehensive summary and analyze each contribution, highlight the strength of the best-performing methods, and discuss the possibility of clinical translations of such methods into the clinic. Our analysis revealed that the participants improved dice coefficient metrics from 0.8607 in 2020 to 0.8993 in 2021 despite adding diverse and challenging frames (containing irregular, smaller, sessile, or flat polyps), which are frequently missed during a routine clinical examination. For the instrument segmentation task, the best team obtained a mean Intersection over union metric of 0.9364. For the transparency task, a multi-disciplinary team, including expert gastroenterologists, accessed each submission and evaluated the team based on open-source practices, failure case analysis, ablation studies, usability and understandability of evaluations to gain a deeper understanding of the models’ credibility for clinical deployment. The best team obtained a final transparency score of 21 out of 25. Through the comprehensive analysis of the challenge, we not only highlight the advancements in polyp and surgical instrument segmentation but also encourage subjective evaluation for building more transparent and understandable AI-based colonoscopy systems. Moreover, we discuss the need for multi-center and out-of-distribution testing to address the current limitations of the methods to reduce the cancer burden and improve patient care. • We present a detailed analysis of the Medico 2020 and MedAI 2021 challenges that are aimed at advancing automated polyp and instrument segmentation in colonoscopy for early colorectal cancer diagnosis by using novel deep learning methods. • To the best of our knowledge, MedAI 2021 is the first challenge to evaluate the transparency in both GI endoscopy and colonoscopy. Through the challenge, we invited the participants to list package dependencies and architecture code (with instructions for building, compiling, and training) and share trained model weights in a standardized format. Additionally, we invited participants to include the code for model evaluation and provide repository licensing information to enable others to use the code and the trained model responsibly. Moreover, we asked the participants to explain model predictions using intermediate heatmaps, perform ablation studies, conduct a thorough failure analysis, and share their code for reproducing the results. Finally, we performed a subjective evaluation by including an expert gastroenterologist in the group and gave the final transparency score based on the usefulness and understandability of the results. Our initiative aims to promote transparency in AI research and foster the development of reliable, interpretable, and trustworthy algorithms for use in medical image segmentation. • We provide a comparative analysis of the 34 proposed methods in both challenges (3 subtasks), covering small details of each team in the form of Tables, qualitative and quantitative results (failure analysis), and an in-depth analysis of the findings. • We explore trust, safety, interpretability, transparency, and generalizability issues and provide future strategies to overcome the current limitations of developed algorithms.
Debesh Jha, Vanshali Sharma, Debapriya Banik, Debayan Bhattacharya, Kaushiki Roy, Steven Alexander Hicks, Nikhil Kumar Tomar, Vajira Thambawita, Adrian Krenzer, Ge-Peng Ji, Sahadev Poudel, George Batchkala, Saruar Alam, Awadelrahman M. A. Ahmed, Quoc-Huy Trinh, Zeshan Khan, Tien-Phat Nguyen, Shruti Shrestha, Sabari Nathan, Jeonghwan Gwak, Ritika Kumari Jha, Zheyuan Zhang 0001, Alexander Schlaefer, Debotosh Bhattacharjee, Manas Kamal Bhuyan, Pradip K. Das, Deng-Ping Fan, Sravanthi Parasa, Sharib Ali, Michael Riegler 0001, Pål Halvorsen, Thomas de Lange, Ulas Bagci
Medical Image Anal.7
2024 Adaptive Smooth Activation Function for Improved Organ Segmentation and Disease Diagnosis
Koushik Biswas, Debesh Jha, Nikhil Kumar Tomar, Meghana Karri, Amit Reza, Gorkem Durak, Alpay Medetalibeyoglu, Matthew Antalek, Yury Velichko, Daniela P. Ladner, Amir Borhani, Ulas Bagci
MICCAI (9)3
2023 DilatedSegNet: A Deep Dilated Segmentation Network for Polyp Segmentation
Nikhil Kumar Tomar, Debesh Jha, Ulas Bagci
MMM (1)1
2023 FANet: A Feedback Attention Network for Improved Biomedical Image Segmentation
abstract
The increase of available large clinical and experimental datasets has contributed to a substantial amount of important contributions in the area of biomedical image analysis. Image segmentation, which is crucial for any quantitative analysis, has especially attracted attention. Recent hardware advancement has led to the success of deep learning approaches. However, although deep learning models are being trained on large datasets, existing methods do not use the information from different learning epochs effectively. In this work, we leverage the information of each training epoch to prune the prediction maps of the subsequent epochs. We propose a novel architecture called feedback attention network (FANet) that unifies the previous epoch mask with the feature map of the current training epoch. The previous epoch mask is then used to provide hard attention to the learned feature maps at different convolutional layers. The network also allows rectifying the predictions in an iterative fashion during the test time. We show that our proposed feedback attention model provides a substantial improvement on most segmentation metrics tested on seven publicly available biomedical imaging datasets demonstrating the effectiveness of FANet. The source code is available at https://github.com/nikhilroxtomar/FANet.
Nikhil Kumar Tomar, Debesh Jha, Michael Riegler 0001, Håvard D. Johansen, Dag Johansen, Jens Rittscher, Pål Halvorsen, Sharib Ali
IEEE Trans. Neural Networks Learn. Syst.1
2022 Video Capsule Endoscopy Classification using Focal Modulation Guided Convolutional Neural Network
abstract
Video capsule endoscopy is a hot topic in computer vision and medicine. Deep learning can have a positive impact on the future of video capsule endoscopy technology. It can improve the anomaly detection rate, reduce physicians' time for screening, and aid in real-world clinical analysis. Computer-Aided diagnosis (CADx) classification system for video capsule endoscopy has shown a great promise for further improvement. For example, detection of cancerous polyp and bleeding can lead to swift medical response and improve the survival rate of the patients. To this end, an automated CADx system must have high throughput and decent accuracy. In this study, we propose FocalConvNet, a focal modulation network integrated with lightweight convolutional layers for the classification of small bowel anatomical landmarks and luminal findings. FocalCon-vNet leverages focal modulation to attain global context and allows global-local spatial interactions throughout the forward pass. Moreover, the convolutional block with its intrinsic induetive/learning bias and capacity to extract hierarchical features allows our FocalConvNet to achieve favourable results with high throughput. We compare our FocalConvNet with other state-of-the-art (SOTA) on Kvasir-Capsule, a large-scale VCE dataset with 44,228 frames with 13 classes of different anomalies. We achieved the weighted F1-score, recall and Matthews correlation coefficient (MCC) of 0.6734, 0.6373 and 0.2974, respectively, outperforming SOTA methodologies. Further, we obtained the highest throughput of 148.02 images/second rate to establish the potential of FocalConvNet in a real-time clinical environ-ment. The code of the proposed FocalConvNet is available at https://github.com/NoviceMAn-prog/FocalConvNet.
Nikhil Kumar Tomar, Ulas Bagci, Debesh Jha
CBMS2
2022 Automatic Polyp Segmentation with Multiple Kernel Dilated Convolution Network
abstract
The detection and removal of precancerous polyps through colonoscopy is the primary technique for the prevention of colorectal cancer worldwide. However, the miss rate of colorectal polyp varies significantly among the endoscopists. It is well known that a computer-aided diagnosis (CAD) system can assist endoscopists in detecting colon polyps and minimize the variation among endoscopists. In this study, we introduce a novel deep learning architecture, named MKDCNet, for automatic polyp segmentation robust to significant changes in polyp data distribution. MKDCNet is simply an encoder-decoder neural network that uses the pre-trained ResNet50 as the encoder and novel multiple kernel dilated convolution (MKDC) block that expands the field of view to learn more robust and heterogeneous representation. Extensive experiments on four publicly available polyp datasets and cell nuclei dataset show that the proposed MKDCNet outperforms the state-of-the-art methods when trained and tested on the same dataset as well when tested on unseen polyp datasets from different distributions. With rich results, we demonstrated the robustness of the proposed architecture. From an efficiency perspective, our algorithm can process at ($\approx 45$) frames per second on RTX 3090 GPU. MKDCNet can be a strong benchmark for building real-time systems for clinical colonoscopies. The code of the proposed MKDCNet is available at https://github.com/nikhilroxtomar/MKDCNet.
Nikhil Kumar Tomar, Ulas Bagci, Debesh Jha
CBMS1
2022 TGANet: Text-Guided Attention for Improved Polyp Segmentation
Nikhil Kumar Tomar, Debesh Jha, Ulas Bagci, Sharib Ali
MICCAI (3)1
2021 NanoNet: Real-Time Polyp Segmentation in Video Capsule Endoscopy and Colonoscopy
abstract
Deep learning in gastrointestinal endoscopy can assist to improve clinical performance and be helpful to assess lesions more accurately. To this extent, semantic segmentation methods that can perform automated real-time delineation of a region-of-interest, e.g., boundary identification of cancer or pre-cancerous lesions, can benefit both diagnosis and interventions. However, accurate and real-time segmentation of endoscopic images is extremely challenging due to its high operator dependence and high-definition image quality. To utilize automated methods in clinical settings, it is crucial to design lightweight models with low latency such that they can be integrated with low-end endoscope hardware devices. In this work, we propose NanoNet, a novel architecture for the segmentation of video capsule endoscopy and colonoscopy images. Our proposed architecture allows real-time performance and has higher segmentation accuracy compared to other more complex ones. We use video capsule endoscopy and standard colonoscopy datasets with polyps, and a dataset consisting of endoscopy biopsies and surgical instruments, to evaluate the effectiveness of our approach. Our experiments demonstrate the increased performance of our architecture in terms of a trade-off between model complexity, speed, model parameters, and metric performances. Moreover, the resulting models' size is relatively tiny, with only nearly 36,000 parameters compared to traditional deep learning approaches having millions of parameters.
Debesh Jha, Nikhil Kumar Tomar, Sharib Ali, Michael Riegler 0001, Håvard D. Johansen, Dag Johansen, Thomas de Lange, Pål Halvorsen
CBMS2