VLDB 2026 Research / reviewers in the wild / expert
Harshita Sharma
dblp:166/4771
· DBLP profile ↗
18ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0003-4683-2606ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Usability Testing of Instant Messaging Apps Using Measured and Self-Reported DataabstractInstant messaging apps are among the most widely used smartphone apps, but little research has been done on their usability. We conducted an experiment in which thirty volunteers used four common instant messaging apps for 2 to 3 weeks, and we observed their interaction with these apps in controlled conditions. We observed that effectiveness ranged between 69% and 96% and varied significantly (F = 19.551, p < 0.05) among the apps. Efficiency ranged between 50% and 70% and varied significantly (F = 5.447, p < 0.05). There was no major change in effectiveness and efficiency during the intervention period denoting low learnability. The satisfaction of the subjects varied significantly (F = 40.026, p < 0.05) for the apps. The affinity of three of the apps was above 75% denoting that three-fourth of the subjects who used them before the intervention planned to continue using them in the future. There was no strong influence of self-reported performance- and satisfaction related parameters on affinity. Harshita Sharma, Ritu Sibal, Pinaki Chakraborty |
Int. J. Hum. Comput. Interact. | 1 |
| 2025 | Challenges for Responsible AI Design and Workflow Integration in Healthcare: A Case Study of Automatic Feeding Tube Qualification in RadiologyabstractNasogastric tubes (NGTs) are feeding tubes that are inserted through the nose into the stomach to deliver nutrition or medication. If not placed correctly, they can cause serious harm, even death to patients. Recent AI developments demonstrate the feasibility of robustly detecting NGT placement from Chest X-ray images to reduce risks of sub-optimally or critically placed NGTs being missed or delayed in their detection, but gaps remain in clinical practice integration. In this study, we present a human-centered approach to the problem and describe insights derived following contextual inquiry and in-depth interviews with 15 clinical stakeholders. The interviews helped understand challenges in existing workflows, and how best to align technical capabilities with user needs and expectations. We discovered the tradeoffs and complexities that need consideration when choosing suitable workflow stages, target users, and design configurations for different AI proposals. We explored how to balance AI benefits and risks for healthcare staff and patients within broader organizational, technical, and medical-legal constraints. We also identified data issues related to edge cases and data biases that affect model training and evaluation; how data documentation practices influence data preparation and labeling; and how to measure relevant AI outcomes reliably in future evaluations. We discuss how our work informs design and development of AI applications that are clinically useful, ethical, and acceptable in real-world healthcare services. Anja Thieme, Abhijith Rajamohan, Benjamin Cooper, Heather Groombridge, Robert Simister, Barney Wong, Nick Woznitza, Mark A. Pinnock, Maria Wetscherek, Cecily Morrison, Hannah Richardson, Fernando Pérez-García, Stephanie L. Hyland, Shruthi Bannur, Daniel C. Castro, Kenza Bouzid, Anton Schwaighofer, Mercy Ranjit, Harshita Sharma, Matthew P. Lungren, Ozan Oktay, Javier Alvarez-Valle, Aditya V. Nori, Steve K. Harris, Joseph Jacob |
ACM Trans. Comput. Hum. Interact. | 19 |
| 2024 | Multimodal Healthcare AI: Identifying and Designing Clinically Relevant Vision-Language Applications for RadiologyabstractRecent advances in AI combine large language models (LLMs) with vision encoders that bring forward unprecedented technical capabilities to leverage for a wide range of healthcare applications. Focusing on the domain of radiology, vision-language models (VLMs) achieve good performance results for tasks such as generating radiology findings based on a patient’s medical image, or answering visual questions (e.g., “Where are the nodules in this chest X-ray?”). However, the clinical utility of potential applications of these capabilities is currently underexplored. We engaged in an iterative, multidisciplinary design process to envision clinically relevant VLM interactions, and co-designed four VLM use concepts: Draft Report Generation, Augmented Report Review, Visual Search and Querying, and Patient Imaging History Highlights. We studied these concepts with 13 radiologists and clinicians who assessed the VLM concepts as valuable, yet articulated many design considerations. Reflecting on our findings, we discuss implications for integrating VLM capabilities in radiology, and for healthcare AI more generally. Nur Yildirim, Hannah Richardson, Maria Wetscherek, Junaid Bajwa, Joseph Jacob, Mark A. Pinnock, Daniel C. Castro, Shruthi Bannur, Stephanie L. Hyland, Pratik Ghosh, Mercy Ranjit, Kenza Bouzid, Anton Schwaighofer, Fernando Pérez-García, Harshita Sharma, Ozan Oktay, Matthew P. Lungren, Javier Alvarez-Valle, Aditya V. Nori, Anja Thieme |
CHI | 16 |
| 2024 | RadEdit: Stress-Testing Biomedical Vision Models via Diffusion Image Editing
Fernando Pérez-García, Sam Bond-Taylor, Pedro P. Sanchez, Boris van Breugel, Daniel C. Castro, Harshita Sharma, Valentina Salvatelli, Maria Wetscherek, Hannah Richardson, Matthew P. Lungren, Aditya V. Nori, Javier Alvarez-Valle, Ozan Oktay, Maximilian Ilse |
ECCV (12) | 6 |
| 2023 | Learning to Exploit Temporal Structure for Biomedical Vision-Language ProcessingabstractSelf-supervised learning in vision-language processing (VLP) exploits semantic alignment between imaging and text modalities. Prior work in biomedical VLP has mostly relied on the alignment of single image and report pairs even though clinical notes commonly refer to prior images. This does not only introduce poor alignment between the modalities but also a missed opportunity to exploit rich self-supervision through existing temporal content in the data. In this work, we explicitly account for prior images and reports when available during both training and fine-tuning. Our approach, named BioViL-T, uses a CNN-Transformer hybrid multi-image encoder trained jointly with a text model. It is designed to be versatile to arising challenges such as pose variations and missing input images across time. The resulting model excels on downstream tasks both in single- and multi-image setups, achieving state-of-the-art (SOTA) performance on (I) progression classification, (II) phrase grounding, and (III) report generation, whilst offering consistent improvements on disease classification and sentence-similarity tasks. We release a novel multi-modal temporal benchmark dataset, MS-CXR-T, to quantify the quality of vision-language representations in terms of temporal semantics. Our experimental results show the advantages of incorporating prior images and reports to make most use of the data. Shruthi Bannur, Stephanie L. Hyland, Qianchu Liu, Fernando Pérez-García, Maximilian Ilse, Daniel C. Castro, Benedikt Boecking, Harshita Sharma, Kenza Bouzid, Anja Thieme, Anton Schwaighofer, Maria Wetscherek, Matthew P. Lungren, Aditya V. Nori, Javier Alvarez-Valle, Ozan Oktay |
CVPR | 8 |
| 2023 | Exploring the Boundaries of GPT-4 in RadiologyabstractQianchu Liu, Stephanie Hyland, Shruthi Bannur, Kenza Bouzid, Daniel Castro, Maria Wetscherek, Robert Tinn, Harshita Sharma, Fernando Pérez-García, Anton Schwaighofer, Pranav Rajpurkar, Sameer Khanna, Hoifung Poon, Naoto Usuyama, Anja Thieme, Aditya Nori, Matthew Lungren, Ozan Oktay, Javier Alvarez-Valle. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Qianchu Liu, Stephanie L. Hyland, Shruthi Bannur, Kenza Bouzid, Daniel C. Castro, Maria Wetscherek, Robert Tinn, Harshita Sharma, Fernando Pérez-García, Anton Schwaighofer, Pranav Rajpurkar, Sameer Tajdin Khanna, Hoifung Poon, Naoto Usuyama, Anja Thieme, Aditya V. Nori, Matthew P. Lungren, Ozan Oktay, Javier Alvarez-Valle |
EMNLP | 8 |
| 2023 | A Machine Learning Method for Automated Description and Workflow Analysis of First Trimester Ultrasound ScansabstractObstetric ultrasound assessment of fetal anatomy in the first trimester of pregnancy is one of the less explored fields in obstetric sonography because of the paucity of guidelines on anatomical screening and availability of data. This paper, for the first time, examines imaging proficiency and practices of first trimester ultrasound scanning through analysis of full-length ultrasound video scans. Findings from this study provide insights to inform the development of more effective user-machine interfaces, of targeted assistive technologies, as well as improvements in workflow protocols for first trimester scanning. Specifically, this paper presents an automated framework to model operator clinical workflow from full-length routine first-trimester fetal ultrasound scan videos. The 2D+t convolutional neural network-based architecture proposed for video annotation incorporates transfer learning and spatio-temporal (2D+t) modelling to automatically partition an ultrasound video into semantically meaningful temporal segments based on the fetal anatomy detected in the video. The model results in a cross-validation A1 accuracy of 96.10% , F1=0.95 , precision =0.94 and recall =0.95 . Automated semantic partitioning of unlabelled video scans (n=250) achieves a high correlation with expert annotations ( ρ = 0.95, p=0.06 ). Clinical workflow patterns, operator skill and its variability can be derived from the resulting representation using the detected anatomy labels, order, and distribution. It is shown that nuchal translucency (NT) is the toughest standard plane to acquire and most operators struggle to localize high-quality frames. Furthermore, it is found that newly qualified operators spend 25.56% more time on key biometry tasks than experienced operators. Robail Yasrab, Zeyu Fu, He Zhao 0002, Lok Hin Lee, Harshita Sharma, Lior Drukker, Aris T. Papageorghiou, J. Alison Noble |
IEEE Trans. Medical Imaging | 5 |
| 2022 | Visualising Spatio-Temporal Gaze Characteristics for Exploratory Data Analysis in Clinical Fetal Ultrasound ScansabstractVisualising patterns in clinicians' eye movements while interpreting fetal ultrasound imaging videos is challenging. Across and within videos, there are differences in size an d position of Areas-of-Interest (AOIs) due to fetal position, movement and sonographer skill. Currently, AOIs are manually labelled or identified using eye-tracker manufacturer specifications which are not study specific. We propose using unsupervised clustering to identify meaningful AOIs and bi-contour plots to visualise spatio-temporal gaze characteristics. We use Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN) to identify the AOIs, and use their corresponding images to capture granular changes within each AOI. Then we visualise transitions within and between AOIs as read by the sonographer. We compare our method to a standardised eye-tracking manufacturer algorithm. Our method captures granular changes in gaze characteristics which are otherwise not shown. Our method is suitable for exploratory data analysis of eye-tracking data involving multiple participants and AOIs. Clare Teng, Harshita Sharma, Lior Drukker, Aris T. Papageorghiou, J. Alison Noble |
ETRA | 2 |
| 2022 | HAWP: a Dataset for Hindi Arithmetic Word Problem SolvingabstractWord Problem Solving remains a challenging and interesting task in NLP. A lot of research has been carried out to solve different genres of word problems with various complexity levels in recent years. However, most of the publicly available datasets and work has been carried out for English. Recently there has been a surge in this area of word problem solving in Chinese with the creation of large benchmark datastes. Apart from these two languages, labeled benchmark datasets for low resource languages are very scarce. This is the first attempt to address this issue for any Indian Language, especially Hindi. In this paper, we present HAWP (Hindi Arithmetic Word Problems), a dataset consisting of 2336 arithmetic word problems in Hindi. We also developed baseline systems for solving these word problems. We also propose a new evaluation technique for word problem solvers taking equation equivalence into account. Harshita Sharma, Pruthwik Mishra, Dipti Misra Sharma |
LREC | 1 |
| 2022 | Gaze-assisted automatic captioning of fetal ultrasound videos using three-way multi-modal deep neural networksabstractIn this work, we present a novel gaze-assisted natural language processing (NLP)-based video captioning model to describe routine second-trimester fetal ultrasound scan videos in a vocabulary of spoken sonography. The primary novelty of our multi-modal approach is that the learned video captioning model is built using a combination of ultrasound video, tracked gaze and textual transcriptions from speech recordings. The textual captions that describe the spatio-temporal scan video content are learnt from sonographer speech recordings. The generation of captions is assisted by sonographer gaze-tracking information reflecting their visual attention while performing live-imaging and interpreting a frozen image. To evaluate the effect of adding, or withholding, different forms of gaze on the video model, we compare spatio-temporal deep networks trained using three multi-modal configurations, namely: (1) a gaze-less neural network with only text and video as input, (2) a neural network additionally using real sonographer gaze in the form of attention maps, and (3) a neural network using automatically-predicted gaze in the form of saliency maps instead. We assess algorithm performance through established general text-based metrics (BLEU, ROUGE-L, F1 score), a domain-specific metric (ARS), and metrics that consider the richness and efficiency of the generated captions with respect to the scan video. Results show that the proposed gaze-assisted models can generate richer and more diverse captions for clinical fetal ultrasound scan videos than those without gaze at the expense of the perceived sentence structure. The results also show that the generated captions are similar to sonographer speech in terms of discussing the visual content and the scanning actions performed. Mohammad Alsharid, Harshita Sharma, Lior Drukker, Aris T. Papageorghiou, J. Alison Noble |
Medical Image Anal. | 3 |
| 2021 | Knowledge representation and learning of operator clinical workflow from full-length routine fetal ultrasound scan videosabstractUltrasound is a widely used imaging modality, yet it is well-known that scanning can be highly operator-dependent and difficult to perform, which limits its wider use in clinical practice. The literature on understanding what makes clinical sonography hard to learn and how sonography varies in the field is sparse, restricted to small-scale studies on the effectiveness of ultrasound training schemes, the role of ultrasound simulation in training, and the effect of introducing scanning guidelines and standards on diagnostic image quality. The Big Data era, and the recent and rapid emergence of machine learning as a more mainstream large-scale data analysis technique, presents a fresh opportunity to study sonography in the field at scale for the first time. Large-scale analysis of video recordings of full-length routine fetal ultrasound scans offers the potential to characterise differences between the scanning proficiency of experts and trainees that would be tedious and time-consuming to do manually due to the vast amounts of data. Such research would be informative to better understand operator clinical workflow when conducting ultrasound scans to support skills training, optimise scan times, and inform building better user-machine interfaces. This paper is to our knowledge the first to address sonography data science, which we consider in the context of second-trimester fetal sonography screening. Specifically, we present a fully-automatic framework to analyse operator clinical workflow solely from full-length routine second-trimester fetal ultrasound scan videos. An ultrasound video dataset containing more than 200 hours of scan recordings was generated for this study. We developed an original deep learning method to temporally segment the ultrasound video into semantically meaningful segments (the video description). The resulting semantic annotation was then used to depict operator clinical workflow (the knowledge representation). Machine learning was applied to the knowledge representation to characterise operator skills and assess operator variability. For video description, our best-performing deep spatio-temporal network shows favourable results in cross-validation (accuracy: 91.7%), statistical analysis (correlation: 0.98, p < 0.05) and retrospective manual validation (accuracy: 76.4%). For knowledge representation of operator clinical workflow, a three-level abstraction scheme consisting of a Subject-specific Timeline Model (STM), Summary of Timeline Features (STF), and an Operator Graph Model (OGM), was introduced that led to a significant decrease in dimensionality and computational complexity compared to raw video data. The workflow representations were learnt to discriminate between operator skills, where a proposed convolutional neural network-based model showed most promising performance (cross-validation accuracy: 98.5%, accuracy on unseen operators: 76.9%). These were further used to derive operator-specific scanning signatures and operator variability in terms of type, order and time distribution of constituent tasks. Harshita Sharma, Lior Drukker, Pierre Chatelain, Richard Droste, Aris T. Papageorghiou, J. Alison Noble |
Medical Image Anal. | 1 |
| 2020 | Spatio-temporal visual attention modelling of standard biometry plane-finding navigationabstractWe present a novel multi-task neural network called Temporal SonoEyeNet (TSEN) with a primary task to describe the visual navigation process of sonographers by learning to generate visual attention maps of ultrasound images around standard biometry planes of the fetal abdomen, head (trans-ventricular plane) and femur. TSEN has three components: a feature extractor, a temporal attention module (TAM), and an auxiliary video classification module (VCM). A soft dynamic time warping (sDTW) loss function is used to improve visual attention modelling. Variants of the model are trained on a dataset of 280 video clips, each containing one of the three biometry planes and lasting 3-7 seconds, with corresponding real-time recorded gaze tracking data of an experienced sonographer. We report the performances of the different variants of TSEN for visual attention prediction at standard biometry plane detection. The best model performance is achieved using bi-directional convolutional long-short term memory (biCLSTM) in both TAM and VCM, and it outperforms a previous spatial model on all static and dynamic saliency metrics. As an auxiliary task to validate the clinical relevance of the visual attention modelling, the predicted visual attention maps were used to guide standard biometry plane detection in consecutive US video frames. All spatio-temporal TSEN models achieve higher scores compared to a spatial-only baseline; the best performing TSEN model achieves F1 scores on these standard biometry planes of 83.7%, 89.9% and 81.1%, respectively. Richard Droste, Harshita Sharma, Pierre Chatelain, Lior Drukker, Aris T. Papageorghiou, J. Alison Noble |
Medical Image Anal. | 3 |
| 2020 | Evaluation of Gaze Tracking Calibration for Longitudinal Biomedical Imaging StudiesabstractGaze tracking is a promising technology for studying the visual perception of clinicians during image-based medical exams. It could be used in longitudinal studies to analyze their perceptive process, explore human-machine interactions, and develop innovative computer-aided imaging systems. However, using a remote eye tracker in an unconstrained environment and over time periods of weeks requires a certain guarantee of performance to ensure that collected gaze data are fit for purpose. We report the results of evaluating eye tracking calibration for longitudinal studies. First, we tested the performance of an eye tracker on a cohort of 13 users over a period of one month. For each participant, the eye tracker was calibrated during the first session. The participants were asked to sit in front of a monitor equipped with the eye tracker, but their position was not constrained. Second, we tested the performance of the eye tracker on sonographers positioned in front of a cart-based ultrasound scanner. Experimental results show a decrease of accuracy between calibration and later testing of 0.30° and a further degradation over time at a rate of 0.13°. month-1. The overall median accuracy was 1.00° (50.9 pixels) and the overall median precision was 0.16° (8.3 pixels). The results from the ultrasonography setting show a decrease of accuracy of 0.16° between calibration and later testing. This slow degradation of gaze tracking accuracy could impact the data quality in long-term studies. Therefore, the results we present here can help in planning such long-term gaze tracking studies. Pierre Chatelain, Harshita Sharma, Lior Drukker, Aris T. Papageorghiou, J. Alison Noble |
IEEE Trans. Cybern. | 2 |
| 2019 | Captioning Ultrasound Images Automatically
Mohammad Alsharid, Harshita Sharma, Lior Drukker, Pierre Chatelain, Aris T. Papageorghiou, J. Alison Noble |
MICCAI (4) | 2 |
| 2019 | Efficient Ultrasound Image Analysis Models with Sonographer Gaze Assisted Distillation
Arijit Patra, Pierre Chatelain, Harshita Sharma, Lior Drukker, Aris T. Papageorghiou, J. Alison Noble |
MICCAI (4) | 4 |
| 2018 | Multi-task SonoEyeNet: Detection of Fetal Standardized Planes Assisted by Generated Sonographer Attention Maps
Harshita Sharma, Pierre Chatelain, J. Alison Noble |
MICCAI (1) | 2 |
| 2017 | A Comparative Study of Cell Nuclei Attributed Relational Graphs for Knowledge Description and Categorization in Histopathological Gastric Cancer Whole Slide ImagesabstractIn this paper, cell nuclei attributed relational graphs are extensively studied and comparatively analyzed for effective knowledge description and classification in H&E stained whole slide images of gastric cancer. This includes design and implementation of multiple graph variations with diverse tissue component characteristics and architectural properties to obtain enhanced image representations, followed by hierarchical ensemble learning and classification. A detailed comparative analysis of the proposed graph-based methods, also with the established low-level, object-level and high-level image descriptions is performed, that further leads to a hybrid approach combining salient visual information. Quantitative evaluation of investigated methods suggests the suitability of particular graph variants for automatic classification using H&E stained histopathological gastric cancer whole slide images based on HER2 immunohistochemistry. Harshita Sharma, Norman Zerbe, Christine Boger, Stephan Wienert, Olaf Hellwich, Peter Hufnagl |
CBMS | 1 |
| 2015 | Appearance-based necrosis detection using textural features and SVM with discriminative thresholding in histopathological whole slide imagesabstractAutomatic detection of necrosis in histological images is an interesting problem of digital pathology that needs to be addressed. Determination of presence and extent of necrosis can provide useful information for disease diagnosis and prognosis, and the detected necrotic regions can also be excluded before analyzing the remaining living tissue. This paper describes a novel appearance-based method to detect tumor necrosis in histopathogical whole slide images. Studies are performed on heterogeneous microscopic images of gastric cancer containing tissue regions with variation in malignancy level and stain intensity. Textural image features are extracted from image patches to efficiently represent necrotic appearance in the tissue and machine learning is performed using support vector machines followed by discriminative thresholding for our complex datasets. The classification results are quantitatively evaluated for different image patch sizes using two cross validation approaches namely three-fold and leave one out cross validation, and the best average cross validation rate of 85.31% is achieved for the most suitable patch size. Therefore, the proposed method is a promising tool to detect necrosis in heterogeneous whole slide images, showing its robustness to varying visual appearances. Harshita Sharma, Norman Zerbe, Iris Klempert, Sebastian Lohmann, Björn Lindequist, Olaf Hellwich, Peter Hufnagl |
BIBE | 1 |