Neha Gour

dblp:217/5088 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0002-0860-1260ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security Inspection
abstract
Advancements in Computer-Aided Screening (CAS) systems are essential for improving the detection of security threats in X-ray baggage scans. However, current datasets are limited in representing real-world, sophisticated threats and concealment tactics, and existing approaches are constrained by a closed-set paradigm with predefined labels. To address these challenges, we introduce STCray, the first multimodal X-ray baggage security dataset, comprising 46,642 image-caption paired scans across 21 threat categories, generated using an X-ray scanner for airport security. STCray is meticulously developed with our specialized protocol that ensures domain-aware, coherent captions, that lead to the multi-modal instruction following data in X-ray baggage security. This allows us to train a domain-aware visual AI assistant named STING-BEE that supports a range of vision-language tasks, including scene comprehension, referring threat localization, visual grounding, and visual question answering (VQA), establishing novel baselines for multi-modal learning in X-ray baggage security. Further, STING-BEE shows state-of-the-art generalization in cross-domain settings. Code, data, and models are available at https://divs1159.github.io/STING-BEE/.
Divya Velayudhan, Abdelfatah Hassan Ahmed, Mohamad Alansari, Neha Gour, Abderaouf Behouch, Taimur Hassan, Syed Talal Wasim, Nabil Maalej, Muzammal Naseer, Juergen Gall, Mohammed Bennamoun, Ernesto Damiani, Naoufel Werghi
CVPR4
2024 An interpretable dual attention network for diabetic retinopathy grading: IDANet
Amit Bhati, Neha Gour, Pritee Khanna, Aparajita Ojha, Naoufel Werghi
Artif. Intell. Medicine2
2024 Autonomous Localization of X-Ray Baggage Threats via Weakly Supervised Learning
abstract
Autonomous X-ray baggage security screening has shown significant strides recently, proving itself a viable solution to the flaws in manual screening, thanks to advancements in deep learning. However, these data-hungry techniques feed on extensively annotated data involving strenuous labor, impeding their advances in baggage screening. Consequently, we present a context-aware transformer for weakly supervised localization to relieve the annotation burden and provide visual interpretability that aids screeners in threat recognition and researchers in identifying the pitfalls of existing systems. The proposed approach can generalize and localize different types of contraband with only cost-effective binary labels without explicit training on item detection. Context extraction block, integrated into the dual-token framework, generates threat-aware context maps, while the token scoring block focuses on minimizing partial activations. Experimental results surpass state of the art (SOTA) methods in terms of classification and localization accuracies. Furthermore, we analyze failures to determine current vulnerabilities and provide new insights for future research.
Divya Velayudhan, Abdelfatah Hassan Ahmed, Taimur Hassan, Neha Gour, Muhammad Owais, Mohammed Bennamoun, Ernesto Damiani, Naoufel Werghi
IEEE Trans. Ind. Informatics4
2023 A Shallow U-Net with Split-Fused Attention Mechanism for Retinal Vessel Segmentation
abstract
Extraction of retinal vascular parts is an important task in retinal disease diagnosis. Precise segmentation of the retinal vascular pattern is challenging due to its complex structure, overlapping with other anatomical structures, and crucial thin vascular structures. In recent years, complex and heavy deep learning networks have been proposed to segment retinal blood vessels accurately. However, these methods fail to detect the thin vascular structure among different patterns of thick vessels. An attention-based novel architecture is proposed to segment the thin vasculature to address this limitation. The proposed model comprises a shallow U-Net based encoder-decoder architecture with split-fuse attention (SFA) block. The proposed SFA block enables the network to identify the placement of pixels for the tree-shaped vessel patterns at their relative position during the reconstruction phase in the decoder. The attention block aggregates low-level and high-level semantic information, improving the vessel segmentation performance. Experimentation performed on publicly available fundus datasets, DRIVE, HRF, CHASE-DB1, and STARE show that the proposed method performs better than the current state-of-the-art methods. The results demonstrate the adaptability of the proposed model for clinical applications due to its low memory footprint and better performance.
Amit Bhati, Samir Jain, Neha Gour, Pritee Khanna, Aparajita Ojha, Naoufel Werghi
ICIP3
2023 Facet-Level Segmentation of 3d Textures on Cultural Heritage Objects
abstract
Textures in 3D meshes exhibit intrinsic surface variations and are indispensable for various applications, such as retrieval, segmentation, and classification of sculptures, artifacts, and paintings. A 3D texture pattern is a locally repeated surface variation independent of the overall surface geometry and can be determined using the local neighborhood and its characteristics. Texture analysis typically employs computer vision techniques that analyze the entire 3D mesh, derive hand-crafted features, and then utilize the derived features for retrieval or classification. Several traditional and learning-based techniques exist in the literature on surface variations; however, textures are the subject of limited works. We propose a binary classification framework at the facet level for classifying texture and non-texture regions on 3D surfaces. An image sequence is generated at each facet, which serves as input to a deep vision transformer. To generate images at each facet, we construct a grid where each cell is filled with the geometric properties of its neighboring facets. We evaluated the proposed method using two datasets with diverse texture patterns, and the results are encouraging.
Iyyakutti Iyappan Ganapathi, Syed Sadaf Ali, Muhammad Owais, Neha Gour, Sajid Javed, Naoufel Werghi
ICIP4
2023 Strengthening Deep Learning Model for Robust Screening of Volumetric Chest Radiographic Scans
abstract
The emerging deep learning algorithms have shown significant potential in the development of efficient computer-aided diagnosis tools for automated detection of lung infections using chest radiographs. However, many existing methods are slice-based and require manual selection of appropriate slices from the entire CT scan, which is tedious and requires expert radiologists. To overcome these limitations, we propose a recurrent 3D Inception network (R3DI-Net) that sequentially exploits spatial and 3D structural features of the entire CT scan, ultimately leading to improved diagnostic performance. Additionally, the proposed method flexibly handles input CT scans with a variable number of slices without incurring performance degradation. A quantitative evaluation of R3DI-Net was made using a combined collection of three publicly accessible datasets containing a sufficient number of data samples. Our method outperforms various existing methods by achieving remarkable performances of 98.39%, 98.36%, 98.1%, and 98.64% in terms of accuracy, F1-score, sensitivity, and average precision, respectively.
Muhammad Owais, Taimur Hassan, Neha Gour, Iyyakutti Iyappan Ganapathi, Naoufel Werghi
ICIP3
2023 Challenges for ocular disease identification in the era of artificial intelligence
Neha Gour, Muhammad Tanveer 0001, Pritee Khanna
Neural Comput. Appl.1
2022 Ocular diseases classification using a lightweight CNN and class weight balancing on OCT images
Neha Gour, Pritee Khanna
Multim. Tools Appl.1
2020 Speckle denoising in optical coherence tomography images using residual deep convolutional neural network
Neha Gour, Pritee Khanna
Multim. Tools Appl.1
2020 Automated glaucoma detection using GIST and pyramid histogram of oriented gradients (PHOG) descriptors
Neha Gour, Pritee Khanna
Pattern Recognit. Lett.1