VLDB 2026 Research / reviewers in the wild / expert
Sameer K. Antani
dblp:01/1603
· DBLP profile ↗
113ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0002-0040-1387ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 75 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 71 · 6 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 42 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 12 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating hallucinations in synthesized clinical texts to improve multimodal deep learning for dermatologyabstract• An investigation into the effects of pairing synthesized clinical notes with image data to train a multimodal AI algorithm, using dermatology as example problem domain. • Leveraging metadata information to drive clinical note synthesis reduces hallucinations in Large Language Model (LLM) outputs. • Clinical notes generated by different LLMs using metadata lead to similar performance on downstream tasks when paired with real dermatology images. • Combination of multimodal data improves generalization performance on external datasets. Despite recent advancements in the development of foundation models and multimodal (MM) architectures in dermatology, their translation to clinical practice remains limited by the scarcity of large-scale multimodal (MM) datasets, as most publicly available resources are small, unimodal, and lack expressive clinical text. This paper investigates strategies to synthesize and exploit clinical notes paired with dermatological images to effectively train a MM architecture, focusing on solutions to limit the inherently hallucinated contents introduced by Large Language Models (LLMs) and to identify conditions under which synthetic clinical notes can be reliably leveraged. The paper proposes a MM architecture trained on real dermatological images paired with LLM-synthesized clinical notes. We systematically evaluate different note generation strategies, including metadata-guided prompting, alignment of image representations with specific keywords, sentence-level filtering of clinical notes, network architectural designs. Experiments involve 16,000 image-note couples collected from six public datasets for model training and over 37,000 images from fifteen public datasets as external data for generalization assessment. Performance is assessed on cross-modal retrieval and zero-shot learning tasks to quantify robustness and generalization. Results show that metadata inclusion into the prompts reduces the hallucinations within LLM outputs, providing more reliable notes. The resulting MM model trained with these notes show superior performance on multiple downstream tasks. Synthesized clinical notes can be paired with real dermatology images under specific conditions, providing a valuable resource to develop foundation models that can help reduce the dermatologists’ workload. Niccolò Marini, Zhaohui Liang, Sivaramakrishnan Rajaraman, Zhiyun Xue, Sameer K. Antani |
J. Biomed. Informatics | 5 |
| 2025 | The Hidden Threat of Hallucinations in Binary Chest X-Ray Pneumonia ClassificationabstractHallucination in deep learning (DL) classification, where DL models yield confidently erroneous predictions remains a pressing concern. This study investigates whether binary classifiers are truly learning disease-specific features when distinguishing overlapping radiological presentations among pneumonia subtypes on chest X-ray (CXR) images. Specifically, we evaluate if uncertainty measure is a valuable tool in classifying signs of different pathogen-specific subtypes of pneumonia. We evaluated two binary classifiers to classify bacterial pneumonia and viral pneumonia, respectively, from normal CXRs. A third classifier explored the ability to distinguish bacterial from viral pneumonia presentation to highlight our concern regarding the observed hallucinations in the former cases. Our comprehensive analysis computes the Matthews Correlation Coefficient and prediction entropy metrics on a pediatric CXR dataset and reveals that the normal/bacterial and normal/viral classifiers consistently and confidently misclassify the unseen pneumonia subtype to their respective disease class. These findings expose a critical limitation concerning the tendency of binary classifiers to hallucinate by relying on general pneumonia indicators rather than pathogen-specific patterns, thereby challenging their utility in clinical workflows. Sivaramakrishnan Rajaraman, Zhaohui Liang, Niccolò Marini, Zhiyun Xue, Sameer K. Antani |
CBMS | 5 |
| 2024 | Addressing Class Imbalance with Latent Diffusion-based Data Augmentation for Improving Disease Classification in Pediatric Chest X-raysabstractDeep learning (DL) has transformed medical image classification; however, its efficacy is often limited by significant data imbalance due to far fewer cases (minority class) compared to controls (majority class). It has been shown that synthetic image augmentation techniques can simulate clinical variability, leading to enhanced model performance. We hypothesize that they could also mitigate the challenge of data imbalance, thereby addressing overfitting to the majority class and enhancing generalization. Recently, latent diffusion models (LDMs) have shown promise in synthesizing high-quality medical images. This study evaluates the effectiveness of a text-guided image-to-image LDM in synthesizing disease-positive chest X-rays (CXRs) and augmenting a pediatric CXR dataset to improve classification performance. We first establish baseline performance by fine-tuning an ImageNet-pretrained Inception-V3 model on class-imbalanced data for two tasks-normal vs. pneumonia and normal vs. bronchopneumonia. Next, we fine-tune individual text-guided image-to-image LDMs to generate CXRs showing signs of pneumonia and bronchopneumonia. The Inception-V3 model is retrained on an updated data set that includes these synthesized images as part of augmented training and validation sets. Classification performance is compared using balanced accuracy, sensitivity, specificity, F-score, Matthews correlation coefficient (MCC), Kappa, and Youden's index against the baseline performance. Results show that the augmentation significantly improves Youden's index (p<0.05) and markedly enhances other metrics, indicating that data augmentation using LDM-synthesized images is an effective strategy for addressing class imbalance in medical image classification. Sivaramakrishnan Rajaraman, Zhaohui Liang, Zhiyun Xue, Sameer K. Antani |
BIBM | 4 |
| 2023 | Emergency Department Wait Time Forecast based on Semantic and Time Series Patterns in COVID-19 PandemicabstractThis study introduces a new ensemble architecture to improve the wait time forecast for healthcare service in the emergency department (ED) of hospital. The new model first used a fine-tuned text embedding model to extract the contextual semantic meaning of patients’ chief complaint from the electronic patient records to estimate the degree of case urgency and combined to a recurrent neural network to process the regular ED wait time patterns. Four text embedding models including the universal sentence encoder with DAN and transformer encoders, the NNLM, and the Swivel were used for semantic analysis. The results show that the new ensemble model can reduce the prediction errors maximumly by 20.0% in mean of absolute error (MAE), 46.0% in mean of squared error (MSE), and 26.6% in root mean squared error (RMSE). A 5-fold cross validation verified that the new model is robust to the ED wait time prediction before and during the COVID-19 pandemic. We conclude that the new model provides an innovative approach to apply semantic analysis of natural language processing to the domain of time series prediction in the healthcare domain. Zhaohui Liang, Zhiyun Xue, Sivaramakrishnan Rajaraman, Jimmy Huang 0001, Sameer K. Antani |
BIBM | 6 |
| 2023 | A Study on Reducing Big Data Image Annotation Burden Through Iterative Expert-In-The-Loop StrategyabstractA key challenge in development of reliable and robust medical imaging machine learning solution is the lack of annotated data. This problem becomes particularly significant when big data sets are used. These pose a burden on the annotators to manually segment regions of interest which is a labor intensive and tedious approach. One solution toward addressing this challenge is to use an iterative expert-in-the-loop approach where models that are initially, albeit weakly, trained on a small expert segmented data set are progressively used to expand the training data. In this work, we explore the viability of this approach through two segmentation experiments. The first is a challenging problem of segmenting the buccal mucosa region from photographs of the mouth for subsequent detection and classification of lesions aimed at an oral cancer prediction application. The other is to segment the lung region in chest X-ray (CXR) images. For simplicity and to focus on discovering viability and any associated shortcomings, we limited our scope to just using an off-the-shelf U-Net algorithm to determine if this approach to training data expansion improved segmentation results. Our findings show that for the buccal mucosa segmentation in oral photographs, the method achieved up to 10% improvement in Dice Similarity Coefficient over three iterations on a blinded manually segmented hold-out test set before the performance plateaued. However, the training data set size almost doubled in size in two iterations. For CXR lung segmentation, we observe slight performance improvement (1% in one iteration) of the method from its initial model which already has a much higher performance (93%). We analyze the performance of the approach for these data sets and comment on the potential of a human expert-in-the-loop method for training data expansion for unlabeled or weakly-labeled medical imaging data. Evanjelin Mahmoodi, Zhiyun Xue, Sivaramakrishnan Rajaraman, Sameer K. Antani |
BIBM | 4 |
| 2023 | Can deep adult lung segmentation models generalize to the pediatric population?abstractLung segmentation in chest X-rays (CXRs) is an important prerequisite for improving the specificity of diagnoses of cardiopulmonary diseases in a clinical decision support system. Current deep learning models for lung segmentation are trained and evaluated on CXR datasets in which the radiographic projections are captured predominantly from the adult population. However, the shape of the lungs is reported to be significantly different across the developmental stages from infancy to adulthood. This might result in age-related data domain shifts that would adversely impact lung segmentation performance when the models trained on the adult population are deployed for pediatric lung segmentation. In this work, our goal is to (i) analyze the generalizability of deep adult lung segmentation models to the pediatric population and (ii) improve performance through a stage-wise, systematic approach consisting of CXR modality-specific weight initializations, stacked ensembles, and an ensemble of stacked ensembles. To evaluate segmentation performance and generalizability, novel evaluation metrics consisting of mean lung contour distance (MLCD) and average hash score (AHS) are proposed in addition to the multi-scale structural similarity index measure (MS-SSIM), the intersection of union (IoU), Dice score, 95% Hausdorff distance (HD95), and average symmetric surface distance (ASSD). Our results showed a significant improvement (p < 0.05) in cross-domain generalization through our approach. This study could serve as a paradigm to analyze the cross-domain generalizability of deep segmentation models for other medical imaging modalities and applications. Sivaramakrishnan Rajaraman, Feng Yang 0010, Ghada Zamzmi, Zhiyun Xue, Sameer K. Antani |
Expert Syst. Appl. | 5 |
| 2023 | Guest Editorial Multimodal Learning in Medical Imaging InformaticsabstractThe papers in this special section focus on multimodal learning in medical imaging informatics. Enormous amounts of health-related data are produced daily, such as those from personal devices, e.g., fitness trackers or mobile applications, ambient sensors, clinical data in electronic health records, pathology reports, lab results, medical images, voice recordings, etc. The practice of modern medicine increasingly relies on data from multiple sources to guide better care. Toward this goal, the papers in this section describe the tools and techniques that integrate multiple data types to describe a particular medical event/case toward developing higher confidence in their decision-making and guidance. KC Santosh, Sameer K. Antani |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | Extraction of Ruler Markings For Estimating Physical Size of Oral LesionsabstractSmall ruler tapes are commonly placed on the surface of the human body as a simple and efficient reference for capturing on images the physical size of a lesion. In this paper, we describe our proposed approach for automatically extracting the measurement information from a ruler in oral cavity images which are taken during oral cancer screening and follow up. The images were taken during a study that aims to investigate the natural history of histologically defined oral cancer precursor lesions and identify epidemiologic factors and molecular markers associated with disease progression. Compared to similar work in the literature proposed for other applications where images are captured with greater consistency and in more controlled situations, we address additional challenges that our application faces in real world use and with analysis of retrospectively collected data. Our approach considers several conditions with respect to ruler style, ruler visibility completeness, and image quality. Further, we provide multiple ways of extracting ruler markings and measurement calculation based on specific conditions. We evaluated the proposed method on two datasets obtained from different sources and examined cross-dataset performance. Zhiyun Xue, Kelly Yu, Paul C. Pearlman, Tseng-Cheng Chen, Chun-Hung Hua, Chung Jan Kang, Chih-Yen Chien, Ming-Hsui Tsai, Cheng-Ping Wang, Anil K. Chaturvedi, Sameer K. Antani |
ICPR | 11 |
| 2022 | Real-time echocardiography image analysis and quantification of cardiac indicesabstractDeep learning has a huge potential to transform echocardiography in clinical practice and point of care ultrasound testing by providing real-time analysis of cardiac structure and function. Automated echocardiography analysis is benefited through use of machine learning for tasks such as image quality assessment, view classification, cardiac region segmentation, and quantification of diagnostic indices. By taking advantage of high-performing deep neural networks, we propose a novel and eicient real-time system for echocardiography analysis and quantification. Our system uses a self-supervised modality-specific representation trained using a publicly available large-scale dataset. The trained representation is used to enhance the learning of target echo tasks with relatively small datasets. We also present a novel Trilateral Attention Network (TaNet) for real-time cardiac region segmentation. The proposed network uses a module for region localization and three lightweight pathways for encoding rich low-level, textural, and high-level features. Feature embeddings from these individual pathways are then aggregated for cardiac region segmentation. This network is fine-tuned using a joint loss function and training strategy. We extensively evaluate the proposed system and its components, which are echo view retrieval, cardiac segmentation, and quantification, using four echocardiography datasets. Our experimental results show a consistent improvement in the performance of echocardiography analysis tasks with enhanced computational eiciency that charts a path toward its adoption in clinical practice. Specifically, our results show superior real-time performance in retrieving good quality echo from individual cardiac view, segmenting cardiac chambers with complex overlaps, and extracting cardiac indices that highly agree with the experts' values. The source code of our implementation can be found in the project's GitHub page. Ghada Zamzmi, Sivaramakrishnan Rajaraman, Li-Yueh Hsu, Vandana Sachdev, Sameer K. Antani |
Medical Image Anal. | 5 |
| 2021 | Semi-Supervised Learning for Cervical Precancer DetectionabstractConvolutional neural networks have become the paradigm of choice for medical image classification applications. Recent research results have demonstrated that deep learning can provide a promising solution for cervical precancer, which is the direct precursor to invasive cervical cancer. However, labeled large datasets are required to develop robust, reliable, and portable deep learning algorithms. This paper presents a study of semi-supervised learning with split-attention models for cervical precancer classification using data derived from two large studies conducted by the U.S. National Cancer Institute. In this work, we examine semi-supervised learning with the ResNeSt50 architecture and observe a significant boost in performance over transfer learning from pre-trained ImageNet weights. We also analyze the issue of specular reflections which is very common in cervical photographic images. Specular reflection occurs as bright spots saturated with white light from the illuminant that occurs due to the presence of moisture and can distract machine learning algorithms. We explore various augmentation techniques to solve specular reflection problems to improve the visual quality of results. As a result, our approach brings significant performance improvements (82.02% accuracy) with potential application in AI device-assisted decision-making. Sandeep Angara, Zhiyun Xue, Sameer K. Antani |
CBMS | 4 |
| 2021 | A Deep Clustering Method For Analyzing Uterine Cervix Images Across Imaging DevicesabstractVisual inspection of the cervix with acetic acid (VIA), though error prone, has long been used for screening women and to guide management for cervical cancer. The automated visual evaluation (AVE) technique, in which deep learning is used to predict precancer based on a digital image of the acetowhitened cervix, has demonstrated its promise as a low-cost method to improve on human performance. However, there are several challenges in moving AVE beyond proof-of-concept and deploying it as a practical adjunct tool in visual screening. One of them is making AVE robust across images captured using different devices. We propose a new deep learning based clustering approach to investigate whether the images taken by three different devices (a common smartphone, a custom smartphone-based handheld device for cervical imaging, and a clinical colposcope equipped with SLR digital camera-based imaging capability) can be well distinguished from each other with respect to the visual appearance/content within their cervix regions. We argue that disparity in visual appearance of a cervix across devices could be a significant confounding factor in training and generalizing AVE performance. Our method consists of four components: cervix region detection, feature extraction, feature encoding, and clustering. Multiple experiments are conducted to demonstrate the effectiveness of each component and compare alternative methods in each component. Our proposed method achieves high clustering accuracy (97%) and significantly outperforms several representative deep clustering methods on our dataset. The high clustering performance indicates the images taken from these three devices are different with respect to visual appearance. Our results and analysis establish a need for developing a method that minimizes such variance among the images acquired from different devices. It also recognizes the need for large number of training images from different sources for robust device-independent AVE performance worldwide. Zhiyun Xue, Kanan T. Desai, Anabik Pal, Olusegun Kayode Ajenifuja, Clement Akinfolarin Adepiti, L. Rodney Long, Mark Schiffman, Sameer K. Antani |
CBMS | 9 |
| 2021 | Life science and its implications for society - (in addition to COVID-19)abstractThis multidisciplinary panel of experts in medicine considers the applications and impacts of technological innovations like Artificial Intelligence, automation, and the Internet of Things, focusing especially on addressing global health challenges, particularly for the post-COVID-19 pandemic era, including in developing nations and underserved populations. Panelists will discuss the opportunities and challenges of telemedicine, cybercare, homecare, treating noncommunicable diseases and preventing communicable diseases, as well as the development of reliable policy and standards for privacy and security of digital innovations. Sameer K. Antani, Luis Kun, Carole Carey, Thenusha Satsoruban, Nahum Gershon, Mohamad Sawan |
ISTAS | 1 |
| 2021 | Selective synthetic augmentation with HistoGAN for improved histopathology image classification
Yuan Xue 0002, Jiarong Ye, Qianying Zhou, L. Rodney Long, Sameer K. Antani, Zhiyun Xue, Carl Cornwell, Richard Zaino, Keith C. Cheng, Sharon X. Huang |
Medical Image Anal. | 5 |
| 2021 | Clustering-Based Dual Deep Learning Architecture for Detecting Red Blood Cells in Malaria Diagnostic SmearsabstractComputer-assisted algorithms have become a mainstay of biomedical applications to improve accuracy and reproducibility of repetitive tasks like manual segmentation and annotation. We propose a novel pipeline for red blood cell detection and counting in thin blood smear microscopy images, named RBCNet, using a dual deep learning architecture. RBCNet consists of a U-Net first stage for cell-cluster or superpixel segmentation, followed by a second refinement stage Faster R-CNN for detecting small cell objects within the connected component clusters. RBCNet uses cell clustering instead of region proposals, which is robust to cell fragmentation, is highly scalable for detecting small objects or fine scale morphological structures in very large images, can be trained using non-overlapping tiles, and during inference is adaptive to the scale of cell-clusters with a low memory footprint. We tested our method on an archived collection of human malaria smears with nearly 200,000 labeled cells across 965 images from 193 patients, acquired in Bangladesh, with each patient contributing five images. Cell detection accuracy using RBCNet was higher than 97 %. The novel dual cascade RBCNet architecture provides more accurate cell detections because the foreground cell-cluster masks from U-Net adaptively guide the detection stage, resulting in a notably higher true positive and lower false alarm rates, compared to traditional and other deep learning methods. The RBCNet pipeline implements a crucial step towards automated malaria diagnosis. Yasmin M. Kassim, Kannappan Palaniappan, Feng Yang 0010, Mahdieh Poostchi, Nila Palaniappan, Richard James Maude, Sameer K. Antani, Stefan Jäger 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2020 | Enhancing Recall Using Data Cleaning for Biomedical Big DataabstractIn clinical practice, large amounts of heterogeneous medical data are generated on a daily basis. This data has the potential to be used for biomedical research and as a diagnostic reference for physicians. However, leveraging heterogeneous data for analysis requires integrating it first. Integration process includes a pre-processing data cleaning phase that eliminates inconsistencies and errors originating from each data source. In this paper, we describe a workflow for cleaning heterogeneous biomedical data sources. Our novel data cleaning approach can be applied for replacement of missing text and to improve the number of relevant cases retrieved by search queries. When the threshold for missing category replacement is met, our results show that our method achieves a missing content replacement precision of 85%, which represents an improvement of 18% over the baseline state of our datasets. Priya Deshpande, Alexander Rasin, Roselyne Tchoua, Jacob D. Furst, Daniela Raicu, Sameer K. Antani |
CBMS | 6 |
| 2020 | Segmentation of Anterior Tissues in Craniofacial Cone-Beam CT ImagesabstractCone-beam computed tomography (CBCT) images are used in craniofacial research for diagnosing dentofacial deformities, skeletal malocclusion severity and to assist in virtual surgical planning. There is a need for automated guidance in predicting regions that could most benefit from surgical intervention. As a part of the effort to conduct such experiments, it is preferable to remove soft tissues in the craniofacial region in CBCT images. However, this front end "data preparation" step is non-trivial for CBCT images due to the inherent fluctuations in the intensity of tissues and bones caused by photon scattering of cone beam shaped X-rays during image acquisition. In this paper, we describe our automated segmentation approach for segmenting anterior tissues in more than 600 3D CBCT images with good result, by combining a selected set of 2D image processing techniques in conjunction with certain facial biometric parameters. Dharitri Misra, Michael Gill, Janice Lee, Sameer K. Antani |
CBMS | 4 |
| 2020 | Synthetic Sample Selection via Reinforcement Learning
Jiarong Ye, Yuan Xue 0002, L. Rodney Long, Sameer K. Antani, Zhiyun Xue, Keith C. Cheng, Sharon X. Huang |
MICCAI (1) | 4 |
| 2020 | Unified deep neural network for segmentation and labeling of multipanel biomedical figuresabstractAbstract Recent efforts in biomedical visual question answering (VQA) research rely on combined information gathered from the image content and surrounding text supporting the figure. Biomedical journals are a rich source of information for such multimodal content indexing. For multipanel figures in these journals, it is critical to develop automatic figure panel splitting and label recognition algorithms to associate individual panels with text metadata in the figure caption and the body of the article. Challenges in this task include large variations in figure panel layout, label location, size, contrast to background, and so on. In this work, we propose a deep convolutional neural network, which splits the panels and recognizes the panel labels in a single step. Visual features are extracted from several layers at various depths of the backbone neural network and organized to form a feature pyramid. These features are fed into classification and regression networks to generate candidates of panels and their labels. These candidates are merged to create the final panel segmentation result through a beam search algorithm. We evaluated the proposed algorithm on the ImageCLEF data set and achieved better performance than the results reported in the literature. In order to thoroughly investigate the proposed algorithm, we also collected and annotated our own data set of 10,642 figures. The experiments, trained on 9,642 figures and evaluated on the remaining 1,000 figures, show that combining panel splitting and panel label recognition mutually benefit each other. Jie Zou 0004, George R. Thoma, Sameer K. Antani |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2020 | Recent trends in image processing and pattern recognition
KC Santosh, Sameer K. Antani |
Multim. Tools Appl. | 2 |
| 2020 | Deep Learning for Smartphone-Based Malaria Parasite Detection in Thick Blood SmearsabstractOBJECTIVE: This work investigates the possibility of automated malaria parasite detection in thick blood smears with smartphones. METHODS: We have developed the first deep learning method that can detect malaria parasites in thick blood smear images and can run on smartphones. Our method consists of two processing steps. First, we apply an intensity-based Iterative Global Minimum Screening (IGMS), which performs a fast screening of a thick smear image to find parasite candidates. Then, a customized Convolutional Neural Network (CNN) classifies each candidate as either parasite or background. Together with this paper, we make a dataset of 1819 thick smear images from 150 patients publicly available to the research community. We used this dataset to train and test our deep learning method, as described in this paper. RESULTS: A patient-level five-fold cross-evaluation demonstrates the effectiveness of the customized CNN model in discriminating between positive (parasitic) and negative image patches in terms of the following performance indicators: accuracy (93.46% ± 0.32%), AUC (98.39% ± 0.18%), sensitivity (92.59% ± 1.27%), specificity (94.33% ± 1.25%), precision (94.25% ± 1.13%), and negative predictive value (92.74% ± 1.09%). High correlation coefficients (>0.98) between automatically detected parasites and ground truth, on both image level and patient level, demonstrate the practicality of our method. CONCLUSION: Promising results are obtained for parasite detection in thick blood smears for a smartphone application using deep learning methods. SIGNIFICANCE: Automated parasite detection running on smartphones is a promising alternative to manual parasite counting for malaria diagnosis, especially in areas lacking experienced parasitologists. Feng Yang 0010, Mahdieh Poostchi, Zhou Zhou 0014, Kamolrat Silamut, Jian Yu 0001, Richard James Maude, Stefan Jäger 0001, Sameer K. Antani |
IEEE J. Biomed. Health Informatics | 9 |
| 2019 | Comparing Deep Learning Models for Multi-cell Classification in Liquid-based Cervical Cytology Images
Sudhir Sornapudi, Gregory T. Brown, Zhiyun Xue, L. Rodney Long, Lisa Allen, Sameer K. Antani |
AMIA | 6 |
| 2019 | Optic Disc and Cup Segmentation for Glaucoma Characterization Using Deep LearningabstractGlaucoma is one of the most common eye diseases that can cause irreversible vision loss due to damage to the optic nerve. Ophthalmologists consider a cup to optic disc ratio greater than 0.3 to be suggestive of glaucoma. Unfortunately, there is high variability among ophthalmologists in estimating the ratio since it is not easy to reliably measure optic disc and cup areas in a fundus image. Therefore, this paper proposes automatic methods to segment the optic disc and cup areas. There are two steps to estimate the ratio: region of interest (ROI) area detection (where optic disc is in the center) from a fundus image, followed by optic disc and cup segmentation. This paper focuses on automated methods to segment the optic disc and cup from the ROI. Fully convolutional networks (FCN) with U-Net architectures are used for the segmentation. The RIGA dataset (composed of three different fundus image datasets: MESSIDOR, Bin Rushed, and Magrabi), containing 750 fundus images, is used to train and test the FCNs. Our proposed FCNs show relatively better performance than other existing algorithms. The best segmentation results for optic disc show 0.95 Jaccard index, 0.98 F-measure, and 0.99 accuracy. The best segmentation results for cup show 0.80 Jaccard index, 0.88 F-measure, and 0.99 accuracy. Loc Q. Tran, Emily Y. Chew, Sameer K. Antani |
CBMS | 4 |
| 2019 | Echo Doppler Flow Classification and Goodness Assessment with Convolutional Neural NetworksabstractDoppler Echocardiography is critical for measuring abnormal cardiac function and diagnosing valvular stenosis and regurgitation. The current practice for assessing and interpreting Doppler echo images is time-consuming and depends highly on the experience of the operator. The limitations of this practice can be mitigated using fully automated intelligent systems. Essential first steps toward comprehensive computer-assisted Doppler echocardiographic interpretation include automatic classification into view/flow categories and goodness assessment of these flows. In this paper, we propose a deep learning-based method for Doppler flow classification and goodness assessment. The method has been trained on labeled images representing a wide range of real-world clinical variation. Our method, when evaluated on unseen data, achieved overall accuracies of 91.6% and 88.9% for flow classification and goodness assessment, respectively. While further research is needed, these results are encouraging and prove the feasibility of using fully automated intelligent systems for analyzing and interpreting Doppler echo images. Ghada Zamzmi, Li-Yueh Hsu, Vandana Sachdev, Sameer K. Antani |
ICMLA | 5 |
| 2019 | Synthetic Augmentation and Feature-Based Filtering for Improved Cervical Histopathology Image Classification
Yuan Xue 0002, Qianying Zhou, Jiarong Ye, L. Rodney Long, Sameer K. Antani, Carl Cornwell, Zhiyun Xue, Sharon X. Huang |
MICCAI (1) | 5 |
| 2019 | Guest Editorial Small Things and Big Data: Controversies and Challenges in Digital HealthcareabstractThe papers in this special section focus on the challenges faced in the digital healthcare market. Recent advances in information and communication technologies (ICT), as well as biomedical engineering, sensor technology and data science, have acted as catalysts for significant developments in the sector of health care, strongly affecting medical diagnosis, patient and healthcare management, disease treatment and health education. In fact, small wearable, disposable sensors, implantable devices or medical devices, as well as elementary services are being featured as keys for monitoring health and facilitating well-being. Panagiotis D. Bamidis, Stathis Th. Konstantinidis, Pedro Pereira Rodrigues, Sameer K. Antani, Daniela Giordano |
IEEE J. Biomed. Health Informatics | 4 |
| 2018 | Gender Detection from Spine X-Ray Images Using Deep LearningabstractThe algorithm described in this paper aims to classify the spine x-ray images according to image characteristics that exhibit gender. We developed a customized sequential CNN model which is trained from scratch using the spine images first and tested it on the NHANES II dataset hosted by the U.S. National Library of Medicine (NLM). Aiming to improve the performance, we then developed a method for extracting the region-of-interest (ROI) in the cervical spine images using a content-based image retrieval (CBIR) method and compared the results of using the original images vs. the ROI images. Later, we applied/tested the method of fine-tuning a DenseNet model that was pre-trained with the ImageNet dataset with the spine images, and this approach gets the best result, achieving classification accuracy of 99% for cervical spine image set and 98% for the lumbar spine image set. Zhiyun Xue, Sivaramakrishnan Rajaraman, L. Rodney Long, Sameer K. Antani, George R. Thoma |
CBMS | 4 |
| 2018 | Multimodal Recurrent Model with Attention for Automated Radiology Report Generation
Yuan Xue 0002, Tao Xu 0029, L. Rodney Long, Zhiyun Xue, Sameer K. Antani, George R. Thoma, Sharon X. Huang |
MICCAI (1) | 5 |
| 2018 | Automated Chest X-Ray Screening: Can Lung Region Symmetry Help Detect Pulmonary Abnormalities?abstractOur primary motivator is the need for screening HIV+ populations in resource-constrained regions for exposure to Tuberculosis, using posteroanterior chest radiographs (CXRs). The proposed method is motivated by the observation that radiological examinations routinely conduct bilateral comparisons of the lung field. In addition, the abnormal CXRs tend to exhibit changes in the lung shape, size, and content (textures), and in overall, reflection symmetry between them. We analyze the lung region symmetry using multi-scale shape features, and edge plus texture features. Shape features exploit local and global representation of the lung regions, while edge and texture features take internal content, including spatial arrangements of the structures. For classification, we have performed voting-based combination of three different classifiers: Bayesian network, multilayer perception neural networks, and random forest. We have used three CXR benchmark collections made available by the U.S. National Library of Medicine and the National Institute of Tuberculosis and Respiratory Diseases, India, and have achieved a maximum abnormality detection accuracy (ACC) of 91.00% and area under the ROC curve (AUC) of 0.96. The proposed method outperforms the previously reported methods by more than 5% in ACC and 3% in AUC. KC Santosh, Sameer K. Antani |
IEEE Trans. Medical Imaging | 2 |
| 2017 | A Kernel Support Vector Machine Trained Using Approximate Global and Exhaustive Local SamplingabstractAGEL-SVM is an extension to a kernel Support Vector Machine (SVM) and is designed for distributed computing using Approximate Global Exhaustive Local sampling (AGEL)-SVM. The dual form of SVM is typically solved using sequential minimal optimization (SMO) which iterates very fast if the full kernel matrix can fit in a computer's memory. AGEL-SVM aims to partition the feature space into sub problems such that the kernel matrix per problem can fit in memory by approximating the data outside each partition. AGEL-SVM has similar Cohen's Kappa and accuracy metrics as the underlying SMO implementation. AGEL-SVM's training times greatly decreased when running on a 128 worker MATLAB pool on Amazon's EC2. Predictor evaluation times are also faster due to a reduction in support vectors per partition. Benjamin Bryant, Hamed Sari-Sarraf, L. Rodney Long, Sameer K. Antani |
BDCAT | 4 |
| 2017 | Named entity recognition in functional neuroimaging literatureabstractHuman neuroimaging research aims to find mappings between brain activity and broad cognitive states. In particular, Functional Magnetic Resonance Imaging (fMRI) allows collecting information about activity in the brain in a non-invasive way. In this paper, we tackle the task of linking brain activity information from fMRI data with named entities expressed in functional neuroimaging literature. For the automatic extraction of those links, we focus on Named Entity Recognition (NER) and compare different methods to recognize relevant entities from fMRI literature. We selected 15 entity categories to describe cognitive states, anatomical areas, stimuli and responses. To cope with the lack of relevant training data, we proposed rule-based methods relying on noun-phrase detection and filtering. We also developed machine learning methods based on Conditional Random Fields (CRF) with morpho-syntactic and semantic features. We constructed a gold standard corpus to evaluate these different NER methods. A comparison of the obtained F1 scores showed that the proposed approaches significantly outperform three state-of-the-art methods in open and specific domains with a best result of 78.79% F1 score in exact span evaluation and 98.40% F1 in inexact span evaluation. Asma Ben Abacha, Alba Garcia Seco de Herrera, L. Rodney Long, Sameer K. Antani, Dina Demner-Fushman |
BIBM | 5 |
| 2017 | Novel Method for Storyboarding Biomedical Videos for Medical InformaticsabstractWe propose a novel method for developing static storyboard for video clips included with biomedical research literature. The technique uses both visual and audio content in the video to select candidate key frames for the storyboard. From the visual channel, the Intra-frames are extracted using FFmpeg tool. IBM Watson speech-to-text service is used to extract words from the audio channel, from which clinically significant concepts (key concepts) are identified using the U.S. National Library of Medicines Repository for Informed Decision Making (RIDeM) service. These concepts are synchronized with the key frames, from which our algorithm selects relevant frames to highlight in the storyboard. In order to test the system, we first created a reference set through a semiautomatic approach, and measure the system performance with informativeness and fidelity metrics. Results from pilot testing, both subjective visual and quantitative metrics, are promising. It is our goal to conduct a formal user evaluation in the future. Sema Candemir, Sameer K. Antani, Zhiyun Xue, George R. Thoma |
CBMS | 2 |
| 2017 | Localizing and Recognizing Labels for Multi-Panel Figures in Biomedical JournalsabstractMulti-panel figures are common in biomedical journals. Often the subpanels are of different types, e.g. x-ray, microscopy, sketch, etc. Visual information retrieval of such figures can significantly benefit from Panel Label Recognition techniques that index figures for search engines, image content tagging, and correlating with figure (sub)captions. It is a challenging task due to large variation in the label locations, sizes, contrast to background, etc. In this work, we propose a 3-stage recognition algorithm. The first stage is formulated as object detection, where we extract Histograms of Oriented Gradient (HOG) features and train a linear Support Vector Machine (SVM) classifier. Label candidates are detected using sliding windows at different locations and scales. We also trained a convolutional deep neural network (CNN) to remove false positives. The second stage is formulated as image classification. We trained a 50-class RBF SVM classifier and estimate the posterior probabilities of each candidate label. The last stage is formulated as sequence classification. We used a beam search algorithm on the posterior probabilities estimated in the second stage along with a set of label sequence constraints to select an optimal label sequence. The algorithm is trained on 9,642 figures, and evaluated on the remaining 1,000 figures shows that the proposed algorithm achieves good precision and recall. Jie Zou 0004, Sameer K. Antani, George R. Thoma |
ICDAR | 2 |
| 2017 | Line Segment-Based Stitched Multipanel Figure Separation for Effective Biomedical CBIRabstractWe present a novel technique to separate panels from stitched multipanel figures appearing in biomedical research articles. Since such figures may comprise images from different imaging modalities, separating them is a crucial first step for effective biomedical content-based image retrieval (CBIR): multimodal biomedical document classification and/or retrieval, for instance. The method applies local line segment detection based on the gray-level pixel changes. It then applies a line vectorization process that connects prominent broken lines along the panel boundaries while eliminating insignificant line segments within the panels. We validated our fully automatic technique on a set of stitched multipanel biomedical figures extracted from articles within the Open Access subset of PubMed Central® repository, and achieved precision and recall of 87.16% and 83.51%, respectively, in less than 0.461[Formula: see text]s per image, on average. We also reported the recent ImageCLEF 2015 competition results that highlight the usefulness of the proposed work. KC Santosh, Aafaque Aafaque, Sameer K. Antani, George R. Thoma |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2017 | Multi-feature based benchmark for cervical dysplasia classification evaluation
Tao Xu 0029, Han Zhang 0010, Cheng Xin, L. Rodney Long, Zhiyun Xue, Sameer K. Antani, Sharon X. Huang |
Pattern Recognit. | 7 |
| 2016 | CNN-based image analysis for malaria diagnosisabstractMalaria is a major global health threat. The standard way of diagnosing malaria is by visually examining blood smears for parasite-infected red blood cells under the microscope by qualified technicians. This method is inefficient and the diagnosis depends on the experience and the knowledge of the person doing the examination. Automatic image recognition technologies based on machine learning have been applied to malaria blood smears for diagnosis before. However, the practical performance has not been sufficient so far. This study proposes a new and robust machine learning model based on a convolutional neural network (CNN) to automatically classify single cells in thin blood smears on standard microscope slides as either infected or uninfected. In a ten-fold cross-validation based on 27,578 single cell images, the average accuracy of our new 16-layer CNN model is 97.37%. A transfer learning model only achieves 91.99% on the same images. The CNN model shows superiority over the transfer learning model in all performance indicators such as sensitivity (96.99% vs 89.00%), specificity (97.75% vs 94.98%), precision (97.73% vs 95.12%), F1 score (97.36% vs 90.24%), and Matthews correlation coefficient (94.75% vs 85.25%). Zhaohui Liang, Andrew Powell 0001, Ilker Ersoy, Mahdieh Poostchi, Kamolrat Silamut, Kannappan Palaniappan, Md Amir Hossain, Sameer K. Antani, Richard James Maude, Jimmy Huang 0001, Stefan Jäger 0001, George R. Thoma |
BIBM | 9 |
| 2016 | Modality Classification for Searching Figures in Biomedical LiteratureabstractImage modality classification categorizes images according to their type. It is an important module in the Open-iSM multimodal (text+image) search engine that retrieves figures from biomedical articles. It is a hierarchical classification where on the top level the input figures are classified into two general categories: regular images (X-ray, CT, MRI, photographs, etc.) vs. illustration images (cartoon sketch, charts, graphs, etc.). This binary classification task is challenged by the vast diversity of visual material (image type), and the way it is organized (simple or compound figures). We present two methods for this binary classification: (i) Support Vector Machines (SVM) with manually-selected features, including a feature based on semantic concepts, and, (ii) Deep Learning method which avoids the process of feature handcrafting. Both methods were tested and compared on a dataset of 16400 figures. Both methods achieved good performance (above 95% accuracy). The slightly better performance of the feature-based method demonstrates the effectiveness of the features we chose. Zhiyun Xue, Sameer K. Antani, L. Rodney Long, Dina Demner-Fushman, George R. Thoma |
CBMS | 3 |
| 2016 | A Simple and Efficient Arrowhead Detection Technique in Biomedical ImagesabstractIn biomedical documents/publications, medical images tend to be complex by nature and often contain several regions that are annotated using arrows. In this context, an automated arrowhead detection is a critical precursor to region-of-interest (ROI) labeling and image content analysis. To detect arrowheads, in this paper, images are first binarized using fuzzy binarization technique to segment a set of candidates based on connected component (CC) principle. To select arrow candidates, we use convexity defect-based filtering, which is followed by template matching via dynamic time warping (DTW). The DTW similarity score confirms the presence of arrows in the image. Our test results on biomedical images from imageCLEF 2010 collection shows the interest of the technique, and can be compared with previously reported state-of-the-art results. KC Santosh, Naved Alam, Partha Pratim Roy 0001, Laurent Wendling, Sameer K. Antani, George R. Thoma |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2016 | Preparing a collection of radiology examinations for distribution and retrievalabstractOBJECTIVE: Clinical documents made available for secondary use play an increasingly important role in discovery of clinical knowledge, development of research methods, and education. An important step in facilitating secondary use of clinical document collections is easy access to descriptions and samples that represent the content of the collections. This paper presents an approach to developing a collection of radiology examinations, including both the images and radiologist narrative reports, and making them publicly available in a searchable database. MATERIALS AND METHODS: The authors collected 3996 radiology reports from the Indiana Network for Patient Care and 8121 associated images from the hospitals' picture archiving systems. The images and reports were de-identified automatically and then the automatic de-identification was manually verified. The authors coded the key findings of the reports and empirically assessed the benefits of manual coding on retrieval. RESULTS: The automatic de-identification of the narrative was aggressive and achieved 100% precision at the cost of rendering a few findings uninterpretable. Automatic de-identification of images was not quite as perfect. Images for two of 3996 patients (0.05%) showed protected health information. Manual encoding of findings improved retrieval precision. CONCLUSION: Stringent de-identification methods can remove all identifiers from text radiology reports. DICOM de-identification of images does not remove all identifying information and needs special attention to images scanned from film. Adding manual coding to the radiologist narrative reports significantly improved relevancy of the retrieved clinical documents. The de-identified Indiana chest X-ray collection is available for searching and downloading from the National Library of Medicine (http://openi.nlm.nih.gov/). Dina Demner-Fushman, Marc D. Kohli, Marc B. Rosenman, Sonya E. Shooshan, Laritza Rodriguez, Sameer K. Antani, George R. Thoma, Clement J. McDonald |
J. Am. Medical Informatics Assoc. | 6 |
| 2016 | Nuclei-Based Features for Uterine Cervical Cancer Histology Image Analysis With Fusion-Based ClassificationabstractCervical cancer, which has been affecting women worldwide as the second most common cancer, can be cured if detected early and treated well. Routinely, expert pathologists visually examine histology slides for cervix tissue abnormality assessment. In previous research, we investigated an automated, localized, fusion-based approach for classifying squamous epithelium into Normal, CIN1, CIN2, and CIN3 grades of cervical intraepithelial neoplasia (CIN) based on image analysis of 61 digitized histology images. This paper introduces novel acellular and atypical cell concentration features computed from vertical segment partitions of the epithelium region within digitized histology images to quantize the relative increase in nuclei numbers as the CIN grade increases. Based on the CIN grade assessments from two expert pathologists, image-based epithelium classification is investigated with voting fusion of vertical segments using support vector machine and linear discriminant analysis approaches. Leave-one-out is used for the training and testing for CIN classification, achieving an exact grade labeling accuracy as high as 88.5%. Koyel Banerjee, Ronald Joe Stanley, L. Rodney Long, Sameer K. Antani, George R. Thoma, Rosemary Zuna, Shelliane R. Frazier, Randy H. Moss, William V. Stoecker |
IEEE J. Biomed. Health Informatics | 5 |
| 2015 | Foreign object detection in chest X-raysabstractAutomatic analysis of chest X-ray images is one important approach for screening/identifying pulmonary diseases. The existence of foreign objects in the images hinders the performance of such processing. In this paper, we focus on one type of foreign objects that is often shown in the images of a large dataset of chest X-rays we are working on-the buttons on the gown that the patient is wearing. The method we propose involves four major steps: intensity normalization, low contrast image identification and enhancement, segmentation of lung regions, and button object extraction. Based on the characteristics of the button objects, we applied two methods for the step of button object extraction. One was based on the circular Hough transform; the other was based on the Viola-Jones algorithm. We tested and compared both methods using a ground truth dataset containing 505 button objects. The results demonstrate the effectiveness of the proposed method. Zhiyun Xue, Sema Candemir, Sameer K. Antani, L. Rodney Long, Stefan Jäger 0001, Dina Demner-Fushman, George R. Thoma |
BIBM | 3 |
| 2015 | Stitched Multipanel Biomedical Figure SeparationabstractWe present a novel technique to separate subpanels from stitched multipanel figures appearing in biomedical research articles. Since such figures may comprise images from different imaging modalities, separating them is a critical first step for effective biomedical content-based image retrieval (CBIR). The method applies local line segment detection based on the gray-level pixel changes. It then applies a line vectorization process that connects prominent broken lines along the subpanel boundaries while eliminating insignificant line segments within the subpanels. We have validated our fully automatic technique on a subset of stitched multipanel biomedical figures extracted from articles within the Open Access subset of PubMed Central repository, and have achieved precision and recall of 81.22% and 85.08%, respectively. KC Santosh, Sameer K. Antani, George R. Thoma |
CBMS | 2 |
| 2015 | Automatic Pulmonary Abnormality Screening Using Thoracic Edge MapabstractWe present a novel method for screening pulmonary abnormalities using thoracic edge map in PA chest radiograph (CXR) images. Our particular interest is to aid clinical officers in screening HIV+ populations in resource constrained regions for Tuberculosis (TB). Our work is motivated by the observation that abnormal CXRs tend to exhibit corrupted and/or deformed thoracic edge maps. We study histograms of thoracic edges for all possible orientations of gradients in the range [0, 2π) at different numbers of bins and different pyramid levels. We have used two CXR benchmark collections made available by the U.S. National Library of Medicine, and have achieved a maximum abnormality detection accuracy of 85.92% and area under the ROC curve (AUC) of 0.91 at one second per image, on average, which outperforms the reported state-of-the-art. KC Santosh, Szilárd Vajda, Sameer K. Antani, George R. Thoma |
CBMS | 3 |
| 2015 | Chest X-ray Image View ClassificationabstractThe view information of a chest X-ray (CXR), such as frontal or lateral, is valuable in computer aided diagnosis (CAD) of CXRs. For example, it helps for the selection of atlas models for automatic lung segmentation. However, very often, the image header does not provide such information. In this paper, we present a new method for classifying a CXR into two categories: frontal view vs. lateral view. The method consists of three major components: image pre-processing, feature extraction, and classification. The features we selected are image profile, body size ratio, pyramid of histograms of orientation gradients, and our newly developed contour-based shape descriptor. The method was tested on a large (more than 8,200 images) CXR dataset hosted by the National Library of Medicine. The very high classification accuracy (over 99% for 10-fold cross validation) demonstrates the effectiveness of the proposed method. Zhiyun Xue, Daekeun You, Sema Candemir, Stefan Jäger 0001, Sameer K. Antani, L. Rodney Long, George R. Thoma |
CBMS | 5 |
| 2015 | Automatically Detecting Rotation in Chest Radiographs Using Principal Rib-Orientation Measure for Quality ControlabstractWe present a novel method for detecting rotated lungs in chest radiographs for quality control and augmenting automated abnormality detection. The method computes a principal rib-orientation measure using a generalized line histogram technique for quality control, and therefore augmenting automated abnormality detection. To compute the line histogram, we use line seed filters as kernels to convolve with edge images, and extract a set of lines from the posterior rib-cage. After convolving kernels in all possible orientations in the range [0°, 180°), we measure the angle with maximum magnitude in the line histogram. This measure provides an approximation of the principal chest rib-orientation for each lung. A chest radiograph is upright if the difference between the orientation angles of both lungs with respect to the horizontal axis is negligible. We validate our method on sets of normal and abnormal images and argue that rib orientation can be used for rotation detection in chest radiographs as an aid in quality control during image acquisition. It can also be used for training and testing data sets for computer aided diagnosis research, for example. In our experiments, we achieve a maximum accuracy of approximately 90%. KC Santosh, Sema Candemir, Stefan Jäger 0001, Alexandros Karargyris, Sameer K. Antani, George R. Thoma, Les R. Folio |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2015 | RSILC: Rotation- and Scale-Invariant, Line-based Color-aware descriptor
Sema Candemir, Eugene Borovikov, KC Santosh, Sameer K. Antani, George R. Thoma |
Image Vis. Comput. | 4 |
| 2015 | Multimodal Entity Coreference for Cervical Dysplasia DiagnosisabstractCervical cancer is the second most common type of cancer for women. Existing screening programs for cervical cancer, such as Pap Smear, suffer from low sensitivity. Thus, many patients who are ill are not detected in the screening process. Using images of the cervix as an aid in cervical cancer screening has the potential to greatly improve sensitivity, and can be especially useful in resource-poor regions of the world. In this paper, we develop a data-driven computer algorithm for interpreting cervical images based on color and texture. We are able to obtain 74% sensitivity and 90% specificity when differentiating high-grade cervical lesions from low-grade lesions and normal tissue. On the same dataset, using Pap tests alone yields a sensitivity of 37% and specificity of 96%, and using HPV test alone gives a 57% sensitivity and 93% specificity. Furthermore, we develop a comprehensive algorithmic framework based on Multimodal Entity Coreference for combining various tests to perform disease classification and diagnosis. When integrating multiple tests, we adopt information gain and gradient-based approaches for learning the relative weights of different tests. In our evaluation, we present a novel algorithm that integrates cervical images, Pap, HPV, and patient age, which yields 83.21% sensitivity and 94.79% specificity, a statistically significant improvement over using any single source of information alone. Dezhao Song, Sharon X. Huang, Joseph Patruno, Hector Muñoz-Avila, Jeff Heflin, L. Rodney Long, Sameer K. Antani |
IEEE Trans. Medical Imaging | 8 |
| 2014 | Biomedical image segmentation for semantic visual feature extractionabstractBiomedical photographs comprise diverse optically acquired images. Accurate classification into meaningful subclasses is valuable in biomedical image retrieval systems. Conventional visual descriptors are limited in their ability to assign semantic labels to images for meaningful retrieval. In this paper we propose a Markov random field (MRF)-based biomedical image segmentation method to segment images into meaningful regions that can be associated with semantic labels. We focus on several tissue image types and develop two MRF models: (i) for tissue image detection from large photograph collection; and, (ii) for region segmentation and semantic labeling. Experimental results demonstrate that our method can detect tissue images in about 82% precision, and our proposed visual descriptors computed from the segmentation results outperform existing visual descriptors. This latter result can be effectively used in biomedical image retrieval systems for retrieving tissue images. Daekeun You, Sameer K. Antani, Dina Demner-Fushman, George R. Thoma |
BIBM | 2 |
| 2014 | Rotation Detection in Chest Radiographs Based on Generalized Line Histogram of Rib-OrientationsabstractWe present a generalized line histogram technique to compute global rib-orientation for detecting rotated lungs in chest radiographs. We use linear structuring elements, such as line seed filters, as kernels to convolve with edge images, and extract a set of lines from the posterior rib-cage. After convolving kernels in all possible orientations in the range [0, π], we measure the angle for which the line histogram has maximum magnitude. This measure provides a good approximation of the global chest rib-orientation for each lung. A chest radiograph is said to be upright if the difference between the orientation angles of both lungs with respect to the horizontal axis, is negligible. We validate our method on sets of normal and abnormal images and argue that rib orientation can be used for rotation detection in chest radiographs as aid in quality control during image acquisition, and to discard images from training and testing data sets. In our test, we achieve a maximum accuracy of 90%. KC Santosh, Sema Candemir, Stefan Jäger 0001, Les R. Folio, Alexandros Karargyris, Sameer K. Antani, George R. Thoma |
CBMS | 6 |
| 2014 | Body Segment Classification for Visible Human Cross Section SlicesabstractVisible human data has been widely used in various medical research and computer science applications. We present a new application for this data: a method to classify which body segment a transverse cross section image belongs to. The labeling of the data is created with the guidance of an online body cross section tutorial. The visual properties of the images are represented using a variety of feature descriptors. To avoid problems that arise from the large dimensionality of features, feature selection is applied. The multi-class SVM is employed as the classifier. Both the CT scans and the color photographs of cryosections of the whole body (male and female) are used to test the proposed method. The high performance with overall accuracy above 98% on both the 2160 CT dataset and the 1870 cryosectional photos show the method is very promising. Because of its observed effectiveness on visible human data, we will extend our approach to classify figures in biomedical articles. Zhiyun Xue, Sameer K. Antani, L. Rodney Long, Dina Demner-Fushman, George R. Thoma |
CBMS | 2 |
| 2014 | Does Figure-Text Improve Biomedical Article Retrieval? A Pilot StudyabstractGraphical illustrations are often used in biomedical articles to convey statistical results, schematics, etc. They are frequently accompanied with superimposed text annotations, or figure-text. It is generally assumed that this figure-text provides information that complements associated textual metadata, such as the captions, or bibliographic citations for indexing the figures and enhanced retrieval quality. However, to the best of our knowledge nothing in the literature adequately supports the assumption of information gain. In this article, we report the results of a pilot study that compares image retrieval based on figure captions alone to that using figure-text in addition to captions. In the blind study two judges evaluated a set of figures retrieved on the topic of "lung cancer". We find that figure-text could help improve retrieval performance for specific queries. Daekeun You, Sameer K. Antani, Dina Demner-Fushman, George R. Thoma |
CBMS | 2 |
| 2014 | An MRF Model for Biomedical Image SegmentationabstractWe propose a Markov random field (MRF)-based method to segment photographic biomedical images into three image sub-regions, viz., tissue, photo, and background. Segmentation results are then used to extract local and global visual features to separate images with tissue, such as endoscopic images, from general photographs. Daekeun You, Sameer K. Antani, Dina Demner-Fushman, George R. Thoma |
CBMS | 2 |
| 2014 | Scalable Arrow Detection in Biomedical ImagesabstractIn this paper, we present a scalable arrow detection technique for biomedical images to support information retrieval systems under the purview of content-based image retrieval (CBIR) and text information retrieval (TIR). The idea primarily follows the criteria based on the geometric properties of the arrow, where we introduce signatures from key points associated with it. To handle this, images are first binarized via a fuzzy binarization tool and several regions of interest are labeled accordingly. Each region is used to generate signatures and then compared with the theoretical ones to check their similarity. Our validation over biomedical images shows the advantage of the technique over the most prominent state-of-the-art methods. KC Santosh, Laurent Wendling, Sameer K. Antani, George R. Thoma |
ICPR | 3 |
| 2014 | Integrating visual words as bunch of n-grams for effective biomedical image classificationabstractThe Bag-of-Visual-Words (BoVW) has been frequently used in the classification of image data. However, this modeling approach does not take into consideration the spatial relationships of these words, which is important for similarity measurement between images. We have developed a novel technique to incorporate spatial information of visual words based on the n-grams representation. The method encodes regional layout with a 2-gram representation in the local keypoint neighborhood. The region is divided in two zones to capture the relative orientations of pair-wise visual words. In turn, each image is described by an accumulated vector of 2-grams. Then, we compute the Shannon entropy over a random “bunch” of 2-grams to reduce the dimensionality of the feature vector. We discovered that this reduction technique creates a more discriminative feature vector as well as presents a considerable dimensionality reduction of up to 99%. The final representation is a compact and efficient local image descriptor that encodes frequency and arrangement of visual words. The proposed approach was tested by classifying a standard biomedical image dataset into categories defined by image modality and body part. The experimental results demonstrate the importance of contextual relations of visual words. Our proposed approach improved the classification accuracy compared to the traditional BoVW by 6.03%. Glauco Vitor Pedrosa, Sameer K. Antani, Dina Demner-Fushman, L. Rodney Long, Agma J. M. Traina |
WACV | 3 |
| 2014 | Multimodal biomedical image indexing and retrieval using descriptive text and global feature mapping
Matthew S. Simpson, Dina Demner-Fushman, Sameer K. Antani, George R. Thoma |
Inf. Retr. | 3 |
| 2014 | Lung Segmentation in Chest Radiographs Using Anatomical Atlases With Nonrigid RegistrationabstractThe National Library of Medicine (NLM) is developing a digital chest X-ray (CXR) screening system for deployment in resource constrained communities and developing countries worldwide with a focus on early detection of tuberculosis. A critical component in the computer-aided diagnosis of digital CXRs is the automatic detection of the lung regions. In this paper, we present a nonrigid registration-driven robust lung segmentation method using image retrieval-based patient specific adaptive lung models that detects lung boundaries, surpassing state-of-the-art performance. The method consists of three main stages: 1) a content-based image retrieval approach for identifying training images (with masks) most similar to the patient CXR using a partial Radon transform and Bhattacharyya shape similarity measure, 2) creating the initial patient-specific anatomical model of lung shape using SIFT-flow for deformable registration of training masks to the patient CXR, and 3) extracting refined lung boundaries using a graph cuts optimization approach with a customized energy function. Our average accuracy of 95.4% on the public JSRT database is the highest among published results. A similar degree of accuracy of 94.1% and 91.7% on two new CXR datasets from Montgomery County, MD, USA, and India, respectively, demonstrates the robustness of our lung segmentation approach. Sema Candemir, Stefan Jäger 0001, Kannappan Palaniappan, Jonathan P. Musco, Rahul K. Singh, Zhiyun Xue, Alexandros Karargyris, Sameer K. Antani, George R. Thoma, Clement J. McDonald |
IEEE Trans. Medical Imaging | 8 |
| 2014 | Automatic Tuberculosis Screening Using Chest RadiographsabstractTuberculosis is a major health threat in many regions of the world. Opportunistic infections in immunocompromised HIV/AIDS patients and multi-drug-resistant bacterial strains have exacerbated the problem, while diagnosing tuberculosis still remains a challenge. When left undiagnosed and thus untreated, mortality rates of patients with tuberculosis are high. Standard diagnostics still rely on methods developed in the last century. They are slow and often unreliable. In an effort to reduce the burden of the disease, this paper presents our automated approach for detecting tuberculosis in conventional posteroanterior chest radiographs. We first extract the lung region using a graph cut segmentation method. For this lung region, we compute a set of texture and shape features, which enable the X-rays to be classified as normal or abnormal using a binary classifier. We measure the performance of our system on two datasets: a set collected by the tuberculosis control program of our local county's health department in the United States, and a set collected by Shenzhen Hospital, China. The proposed computer-aided diagnostic system for TB screening, which is ready for field deployment, achieves a performance that approaches the performance of human experts. We achieve an area under the ROC curve (AUC) of 87% (78.3% accuracy) for the first set, and an AUC of 90% (84% accuracy) for the second set. For the first set, we compare our system performance with the performance of radiologists. When trying not to miss any positive cases, radiologists achieve an accuracy of about 82% on this set, and their false positive rate is about half of our system's rate. Stefan Jäger 0001, Alexandros Karargyris, Sema Candemir, Les R. Folio, Jenifer Siegelman, Fiona M. Callaghan, Zhiyun Xue, Kannappan Palaniappan, Rahul K. Singh, Sameer K. Antani, George R. Thoma, Yì Xiáng J. Wáng, Pu-Xuan Lu, Clement J. McDonald |
IEEE Trans. Medical Imaging | 10 |
| 2013 | Graphical Figure Classification Using Data Fusion for Integrating Text and Image FeaturesabstractThis paper describes a multimodal (image + text) learning approach for automatically identifying three graphical figure types commonly found in biomedical literature, namely, diagrams, statistical figures and flow charts. The goal is to improve retrieval of figures from biomedical journal articles. In this article, we describe a data fusion approach to combine information from both text and image sources, believed to contain complementary information. Text information about the image is extracted from the figure caption. The data fusion process includes a hybrid of evolutionary algorithm (EA) and Binary Particle Swarm Optimization (BPSO) called method applied to find an optimal subset of extracted image features. Chi-square statistic and information gain metric are used to select the optimal subset of extracted text features, which along with image features are input to Multi-Layer Perceptron Neural Network classifiers, whose outputs are characterized as fuzzy sets to determine the final classification result. Evaluation performed on 1707 figure images extracted from a test subset of Biome Central® journals extracted from U.S. National Library of Medicine's PubMed Central ® repository yielded classification accuracy as high as 96.1%. Beibei Cheng, Ronald Joe Stanley, Sameer K. Antani, George R. Thoma |
ICDAR | 3 |
| 2013 | Image retrieval from scientific publications: Text and image content processing to separate multipanel figuresabstractImages contained in scientific publications are widely considered useful for educational and research purposes, and their accurate indexing is critical for efficient and effective retrieval. Such image retrieval is complicated by the fact that figures in the scientific literature often combine multiple individual subfigures (panels). Multipanel figures are in fact the predominant pattern in certain types of scientific publications. The goal of this work is to automatically segment multipanel figures—a necessary step for automatic semantic indexing and in the development of image retrieval systems targeting the scientific literature. We have developed a method that uses the image content as well as the associated figure caption to: (1) automatically detect panel boundaries; (2) detect panel labels in the images and convert them to text; and (3) detect the labels and textual descriptions of each panel within the captions. Our approach combines the output of image‐content and text‐based processing steps to split the multipanel figures into individual subfigures and assign to each subfigure its corresponding section of the caption. The developed system achieved precision of 81% and recall of 73% on the task of automatic segmentation of multipanel figures. Emilia Apostolova, Daekeun You, Zhiyun Xue, Sameer K. Antani, Dina Demner-Fushman, George R. Thoma |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2012 | Towards the Creation of a Visual Ontology of Biomedical Imaging Entities
Matthew S. Simpson, Daekeun You, Sameer K. Antani, George R. Thoma, Dina Demner-Fushman |
AMIA | 4 |
| 2012 | Window Classification of Brain CT Images in Biomedical Articles
Zhiyun Xue, Sameer K. Antani, L. Rodney Long, Dina Demner-Fushman, George R. Thoma |
AMIA | 2 |
| 2011 | Review of medical image retrieval systems and future directionsabstractThis paper presents a review of online systems for content-based medical image retrieval (CBIR). The objective of this review is to evaluate the capabilities and gaps in these systems and to determine ways of improving relevance of multi-modal (text and image) information retrieval in the iMedline system, being developed at the National Library of Medicine (NLM). Seven medical information retrieval systems: Figuresearch, BioText, GoldMiner, Yale Image Finder, Yottalook, Image Retrieval for Medical Applications (IRMA), and iMedline have been evaluated here using the system of gaps defined in [1]. Not all of these systems take advantage of the visual information contained in biomedical literature as figures and illustrations. However, all attempt to extract metadata about the image from the full-text of the articles and retrieve figures/images in response to a query. iMedline aims to advance the state-of-the-art in multimodal information retrieval by unifying image and text features in computing relevance. We discuss the shortcomings of these current systems and discuss future directions and next steps in iMedline toward context-based medical image retrieval. Payel Ghosh, Sameer K. Antani, L. Rodney Long, George R. Thoma |
CBMS | 2 |
| 2011 | Biomedical CBIR using "bag of keypoints" in a modified inverted indexabstractThis paper presents a “bag of keypoints” based medical image retrieval approach to cope with a large variety of visually different instances under the same category or modality. Keypoint similarities in the codebook are computed using a quadratic similarity measure. The codebook is implemented using a topology preserving Self Organizing Map (SOM) which represents images as sparse feature vectors and an inverted index is created on top of this to facilitate efficient retrieval. In addition, to increase the retrieval effectiveness, query expansion is performed by exploiting the similarities between the keypoints based on analyzing the local neighborhood structure of the SOM generated codebook. The search is thus query-specific and restricted to a sub-space spanned only by the original and expanded keypoints of the query images. A systematic evaluation of retrieval results on a biomedical image collection of 5000 biomedical images of different modalities, body parts, and orientations shows a halving in computation time (efficiency) and 10% to 15% improvement in precision at each recall level (effectiveness) when compared to individual color, texture, edge-related features. Sameer K. Antani, George R. Thoma |
CBMS | 2 |
| 2011 | Spine X-ray image retrieval using partial vertebral boundariesabstractThe anterior osteophyte (AO) is a bony spur on the vertebra and is symptomatic of osteo-arthritis of the spine. We present advances in our research into matching vertebral boundaries based on pathological (severity) and visual similarity. Proposed image retrieval methods are based on partial shape matching (PSM) that use landmarks along sagittal vertebral outlines that are consistent with those used by medical experts. Besides the two PSM methods that are improved algorithms of our previously developed methods, a new PSM method that is based on a simple but effective localized shape feature is proposed. The methods are evaluated and tested on a dataset of 856 segmented vertebrae and their performance is compared using precision-recall and average precision graphs. The best approaches are combined and integrated into our Web-based Spine Pathology & Image Retrieval System (SPIRS). Zhiyun Xue, L. Rodney Long, Sameer K. Antani, George R. Thoma |
CBMS | 3 |
| 2011 | Detecting Figure-Panel Labels in Medical Journal Articles Using MRFabstractWe present a method for figure-panel (subfigure) label detection and recognition in multi-panel figures extracted from biomedical articles. Figures in biomedical articles often comprise several subfigures that are identified by superimposed panel labels ('A', 'B', ...) which are referenced in the figure caption and discussion in the article body. Splitting such multi-panel figures into individual subfigures is a necessary step for improved multimodal biomedical information retrieval. Prior to feature extraction for indexing and retrieval of biomedical figures it is necessary to classify image content in each subfigure by its modality (X-ray, MRI, CT, etc.) and other relevant criteria. Subfigure labels are valuable in associating individual panels with relevant text in captions and discussion. We propose a 4-step panel label detection method based on Markov Random Field (MRF). Experiments on 515 multi-panel figures and analysis of the results show promising results. We present the successes and identify critical challenges. Daekeun You, Sameer K. Antani, Dina Demner-Fushman, Venu Govindaraju, George R. Thoma |
ICDAR | 2 |
| 2011 | A query expansion framework in image retrieval domain based on local and global analysis
Sameer K. Antani, George R. Thoma |
Inf. Process. Manag. | 2 |
| 2011 | A Learning-Based Similarity Fusion and Filtering Approach for Biomedical Image Retrieval Using SVM Classification and Relevance FeedbackabstractThis paper presents a classification-driven biomedical image retrieval framework based on image filtering and similarity fusion by employing supervised learning techniques. In this framework, the probabilistic outputs of a multiclass support vector machine (SVM) classifier as category prediction of query and database images are exploited at first to filter out irrelevant images, thereby reducing the search space for similarity matching. Images are classified at a global level according to their modalities based on different low-level, concept, and keypoint-based features. It is difficult to find a unique feature to compare images effectively for all types of queries. Hence, a query-specific adaptive linear combination of similarity matching approach is proposed by relying on the image classification and feedback information from users. Based on the prediction of a query image category, individual precomputed weights of different features are adjusted online. The prediction of the classifier may be inaccurate in some cases and a user might have a different semantic interpretation about retrieved images. Hence, the weights are finally determined by considering both precision and rank order information of each individual feature representation by considering top retrieved relevant images as judged by the users. As a result, the system can adapt itself to individual searches to produce query-specific results. Experiment is performed in a diverse collection of 5 000 biomedical images of different modalities, body parts, and orientations. It demonstrates the efficiency (about half computation time compared to search on entire collection) and effectiveness (about 10%-15% improvement in precision at each recall level) of the retrieval approach. Sameer K. Antani, George R. Thoma |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2010 | Integrating image and text information for biomedical information retrievalabstractSummary form only given. The search for relevant and actionable information is key to achieving clinical and research goals in biomedicine. Biomedical information exists in different forms: as text and illustrations in journal articles and other documents, in "images" stored in databases, and as patients' cases in electronic health records. In the context of this work an "image" includes not only biomedical images, but also illustrations, charts, graphs, and other visual material appearing in biomedical journals, electronic health records, and other relevant databases. The tutorial will cover methods and techniques to retrieve information from these entities, by moving beyond conventional text-based searching to combining both text and visual features in search queries. The approaches to meeting these objectives use a combination of techniques and tools from the fields of Information Retrieval (IK), Content-Based Image Retrieval (CBIR), and Natural Language Processing (NLP). The tutorial will discuss steps to improve the retrieval of biomedical literature by targeting the text describing the visual content in articles (figures, including illustrations and images), a rich source of information not typically exploited by conventional bibliographic or full-text databases. Taking this a step further we will explore challenges in finding information relevant to a patient's case from the literature and then link it to the patient's health record. The case is first represented in structured form using both text and image features, and then literature and EHR databases can be searched for similar cases. Further, we will discuss steps to automatically find semantically similar images in image databases, which is an important step in differential diagnosis. Automatic image annotation and retrieval steps will be described that use image features and a combination of image and text features. We explore steps toward generating a "visual ontology", i.e., concepts assigned to image patches. Elements from the visual ontology are called "visual keywords" and are used to find images with similar concepts. The tutorial will demonstrate some of these techniques by demonstrating our Image and Text Search Engine (ITSE), a hybrid system combining NLM's Essie text search engine with CEB's image similarity engine. Sameer K. Antani |
CBMS | 1 |
| 2010 | Comparative study of shape retrieval using feature fusion approachesabstractA thorough comparison of shape similarity distance measures for Content-Based Image Retrieval (CBIR) and application of feature normalization and machine learning has received limited attention. This article reports on the comparison of the performance of several shape similarity algorithms and the effect of several feature normalization methods. We also propose a learning-based feature selection and fusion scheme as an approach to bridge the `semantic gap' between low-level image features and high-level human concepts. The methods are tested on a collection of segmented vertebral boundaries extracted from a subset of digitized x-ray images of the spine from the second National Health and Nutrition Examination Survey (NHANES II). In general the experimental results show that proper multi-feature fusion schemes achieve significantly improved retrieval performance. Haiying Guan, Sameer K. Antani, L. Rodney Long, George R. Thoma |
CBMS | 2 |
| 2010 | "Bag of keypoints"-based biomedical image search with affine covariant region detection and correlation-enhanced similarity matchingabstractThis paper presents a “bag of keypoints” based biomed-ical image retrieval approach by detecting affine covariant regions. These regions refers to a set of pixels or interest points which are invariant to affine transformations, as well as occlusion, lighting and intra-class variations. The interest points are described with the Scale-Invariant Feature Transform (SIFT) and vector quantized to build a visual vocabulary of keypoints. By mapping the interest points extracted from one image to the keypoints in the vocabulary, their occurrences are counted and the resulting histogram is called the “bag of keypoints” for that image similar to the “bag of words” based representation of documents in text retrieval. To exploit the correlations between the keypoints in the collection, a global similarity matrix is constructed to be utilized in a distance measure function to compare the query and database images. A systematic evaluation retrieval results on a biomedical image collection demonstrates around 10-15% improvement in precision at different recall levels for the keypoint-based representation along with the correlation-enhanced distance measure when compared to individual color, texture, edge-related features. Sameer K. Antani, George R. Thoma |
CBMS | 2 |
| 2010 | Automatic extraction of mosaic patterns in uterine cervix imagesabstractMosaic vasculature is one crucial visual sign often indicating the existence of abnormality in the underlying cervix tissues. Automatic detection of this vascular pattern in uterine cervix images is a challenging task, especially in a large dataset, due to the factors such as fuzzy boundary, small vessel caliber, and appearance variation. In this paper, we present a supervised-learning based approach to segment the regions encompassing mosaic vasculatures, hoping to overcome these challenges. It is part of an automatic segmentation scheme that is aimed at assisting gynecologists in the study of cervical cancer. The affectivity of the method was tested and evaluated on a set of clinical uterine cervix images that were manually marked and categorized by medical experts. Zhiyun Xue, L. Rodney Long, Sameer K. Antani, George R. Thoma |
CBMS | 3 |
| 2010 | Local and global Gaussian mixture models for hematoxylin and eosin stained histology image segmentationabstractThis paper presents a new algorithm for hematoxylin and eosin (H&E) stained histology image segmentation. With both local and global clustering, Gaussian mixture models (GMMs) are applied sequentially to extract tissue constituents such as nuclei, stroma, and connecting contents from background. Specifically, local GMM is firstly applied to detect nuclei by scanning the input image, which is followed by global GMM to separate other tissue constituents from background. Regular RGB (red, green and blue) color space is employed individually for the local and global GMMs to make use of the H&E staining features. Experiments on a set of cervix histology images show the improved performance of the proposed algorithm when compared with traditional K-means clustering and state-of-art multiphase level set methods. L. Rodney Long, Sameer K. Antani, George R. Thoma |
HIS | 3 |
| 2010 | Optimal embedding for shape indexing in medical image databases
Xiaoning Qian, Hemant D. Tagare, Robert K. Fulbright, L. Rodney Long, Sameer K. Antani |
Medical Image Anal. | 5 |
| 2010 | Interactive publication: The document as a research tool
George R. Thoma, Glenn Ford, Sameer K. Antani, Dina Demner-Fushman, Matthew S. Simpson |
J. Web Semant. | 3 |
| 2009 | Comparative study of spine vertebra shape retrieval using learning-based feature selectionabstractFeature extraction and selection are two important steps for shape retrieval. Given a data set, a set of features which describe the shape property from different aspects are extracted. Our goal is a learning-based methodology to select the features for improving retrieval performance. Our approach uses both global and local feature descriptors. The global shape features include geometric ones (elongation, eccentricity, roughness, and compactness), Fourier descriptors with complex coordinates, Fourier descriptors with Centroid Contour Distance Curve, Coefficients of Fourier Expansion of Bent function, moment invariants, and local shape features that include turn angle and Distance Across the Shape. We propose a learning-based feature selection algorithm as a strategy for optimizing retrieval performance. We provide results from the vertebra shape dataset created from our database containing spine X-rays from the National Health and Nutrition Examination Survey (NHANES II). Finally, we compare the retrieval performances of feature descriptors on ldquowhole shaperdquo and ldquocorner shaperdquo datasets. The experimental results show that various feature descriptors perform differently on different datasets, and that feature selection schemes improve the retrieval performance significantly. Haiying Guan, Sameer K. Antani, L. Rodney Long, George R. Thoma |
CBMS | 2 |
| 2009 | Lessons learned in developing a low-cost high performance medical imaging clusterabstractThis paper explores the usefulness of the Sony PlayStation 3reg(PS3) for medical image processing. Medical image processing often entails dealing with a large number of high resolution images, requiring a large amount of computational power to process. The PS3 is powered by the cell broadband engine, a microprocessor created by IBM, capable of rapid numeric computation with low power requirements that has helped the unit to become a popular gaming unit. The unit can be repurposed as a low-cost high performance computing platform. In order to demonstrate the computational abilities of the PS3, several basic image processing tasks are implemented and compared with desktop PCs, equipped with general-purpose microprocessors. This article describes lessons learned in the process of building some fundamental image processing tasks. The article also describes the architecture of the cell broadband engine and provides information about developing applications given the architecture. The article also provides an introduction to setting up a high-performance image processing environment with a cluster of such relatively inexpensive PS3 units. Kirt D. Lillywhite, Dah-Jye Lee, Sameer K. Antani, Dong Zhang 0002, L. Rodney Long |
CBMS | 3 |
| 2009 | A medical image retrieval framework in correlation enhanced visual concept feature spaceabstractThis paper presents a medical image retrieval framework that uses visual concepts in a feature space employing statistical models built using a probabilistic multi-class support vector machine (SVM). The images are represented using concepts that comprise color and texture patches from local image regions in a multi-dimensional feature space. A major limitation of concept feature representation is that the structural relationship or spatial ordering between concepts are ignored. We present a feature representation scheme as visual concept structure descriptor (VCSD) that overcomes this challenge and captures both the concept frequency similar to a color histogram and the local spatial relationships of the concepts. A probabilistic framework makes the descriptor robust against classification and quantization errors. Evaluation of the proposed image retrieval framework on a biomedical image dataset with different imaging modalities validates its benefits. Sameer K. Antani, George R. Thoma |
CBMS | 2 |
| 2009 | Linear array image analysis for automated detection of human papillomavirusabstractPersistent infections with carcinogenic human Papillomavirus (HPV) are a necessary cause for cervical cancer, which is the fifth most deadly cancer for women worldwide. Approximately 20 million Americans are currently infected with HPV but only a subset will develop cervical cancer. While a negative HPV test indicates a very low risk for cervical cancer, a positive test cannot discriminate between an innocuous transient infection and a prevalent cancer. Additional information such as HPV genotype and HPV viral load is thought to improve the predictive ability of which women will develop cervical cancer. The visual interpretation of hybridization-strip-based HPV genotyping results, however, is heterogeneous and poorly standardized. This has led to work toward the development of a robust automated image analysis package for HPV genotyping strips. Matthew Wilhelm, Brian Nutter, L. Rodney Long, Sameer K. Antani |
CBMS | 4 |
| 2009 | A system for searching uterine cervix images by visual attributesabstractContent-based indexing and retrieval is gaining increasing interest in the medical domain with the growing size of medical image databases. We present here a Web-accessible retrieval system for searching for similar uterine cervix images based on their visual characteristics. The system operates on a subset of a large database created for archiving patient records collected by two key projects in cervical cancer research. It was developed to bridge the ldquogapsrdquo that hold back the practical adoption of most CBIR systems. This collaboration between engineers and gynecological experts promises to provide a new biomedical resource beyond current text-based searching tools. Zhiyun Xue, Sameer K. Antani, L. Rodney Long, George R. Thoma |
CBMS | 2 |
| 2009 | Using Non-Lexical Features to Identify Effective Indexing Terms for Biomedical Illustrations
Matthew S. Simpson, Dina Demner-Fushman, Charles Sneiderman, Sameer K. Antani, George R. Thoma |
EACL | 4 |
| 2009 | CBIR of spine X-ray images on inter-vertebral disc space and shape profiles using feature ranking and voting consensus
Dah-Jye Lee, Sameer K. Antani, Yuchou Chang, Kent Gledhill, L. Rodney Long, Paul Christensen |
Data Knowl. Eng. | 2 |
| 2009 | Using relevance feedback with short-term memory for content-based spine X-ray image retrieval
Xiaoqian Xu, Dah-Jye Lee, Sameer K. Antani, L. Rodney Long, James K. Archibald |
Neurocomputing | 3 |
| 2009 | Automatic Detection of Anatomical Landmarks in Uterine Cervix ImagesabstractThe work focuses on a unique medical repository of digital cervicographic images ("Cervigrams") collected by the National Cancer Institute (NCI) in longitudinal multiyear studies. NCI, together with the National Library of Medicine (NLM), is developing a unique web-accessible database of the digitized cervix images to study the evolution of lesions related to cervical cancer. Tools are needed for automated analysis of the cervigram content to support cancer research. We present a multistage scheme for segmenting and labeling regions of anatomical interest within the cervigrams. In particular, we focus on the extraction of the cervix region and fine detection of the cervix boundary; specular reflection is eliminated as an important preprocessing step; in addition, the entrance to the endocervical canal (the "os"), is detected. Segmentation results are evaluated on three image sets of cervigrams that were manually labeled by NCI experts. Hayit Greenspan, Shiri Gordon, Gali Zimmerman-Moreno, Shelly Lotenberg, Jose Jeronimo, Sameer K. Antani, L. Rodney Long |
IEEE Trans. Medical Imaging | 6 |
| 2008 | Bridging the Gap: Enabling CBIR in Medical ApplicationsabstractContent-based image retrieval (CBIR) for medical images has received a significant research interest over the past decade as a promising approach to address the data management challenges posed by the rapidly increasing volume of medical image data in use. Articles published in the literature detail the benefits and present impressive results to substantiate potential impact of the technology. However, the benefits have yet to make it to mainstream clinical, biomedical research, or educational use. No major commercial software tools are available for use in medical imaging products, although several are available for commercial stock photo collections. CBIR has had some success in isolated instances in applications on limited data sets addressing specialized medical problems and at biomedical research laboratories and hospitals that are tightly coupled with software developers. This article explores some possible causes of this "gap" in the lack of translation of research into widespread biomedical use and provides some directions to alleviate the problem. Sameer K. Antani, L. Rodney Long, George R. Thoma |
CBMS | 1 |
| 2008 | CBIR of Spine X-Ray Images on Inter-Vertebral Disc Space and Shape ProfilesabstractThere is very limited research published in the literature that applies content-based image retrieval (CBIR) techniques to retrieval of digitized spine X-ray images using a combination of inter-vertebral disc space and shape profiles. We present a novel technique to retrieve vertebra pairs that exhibit a specified disc space narrowing (DSN) and inter-vertebral disc shape profile. DSN is characterized using spatial and geometrical features between two adjacent vertebrae. Initial retrieval results are clustered and used to construct a voting committee to retrieve vertebra pairs with the highest DSN similarity. Experimental results show that the proposed algorithm is a promising approach for disc space-based spine X-ray image retrieval. The overall retrieval accuracy validated by a radiologist is 82.25%. Yuchou Chang, Sameer K. Antani, Dah-Jye Lee, Kent Gledhill, L. Rodney Long, Paul Christensen |
CBMS | 2 |
| 2008 | Web-Based Multi-Observer Segmentation Evaluation ToolabstractMulti-observer segmentation evaluation is useful in the imaging community. We have developed web-based software for automatic performance evaluation of multiple image segmentations which is based on the Baysian decision framework. It computes a probabilistic estimate of the true segmentation (ground truth map) and performance measures for the individual segmentations (sensitivity and specificity). The strength of the tool is that it integrates the two kinds of prior knowledge of segmentations: the truth prior (the prior probability) and the observer prior (the performance measures of observers), which can generate more accurate evaluations. Yaoyao Zhu, Sharon X. Huang, Daniel P. Lopresti, L. Rodney Long, Sameer K. Antani, Zhiyun Xue, George R. Thoma |
CBMS | 5 |
| 2008 | Cervicographic image retrieval by spatial similarity of lesionsabstractThe National Library of Medicine has been developing CervigramFinder, a Web-accessible prototype content-based image retrieval (CBIR) system for cervical cancer research, to retrieve cervicographic images from a large collection with respect to visual characteristics of lesion regions. This paper describes current work on retrieving the images based on similarity of spatial location of lesions. The proposed two-level method takes into account the visual characteristics of cervix lesions, as well as spatial information of shape, size, orientation, and distance. The proposed method was evaluated on a data set of 1000 cervicographic images where multiple lesion boundaries as well as associated location information were marked by medical experts. The simplicity and effectiveness of the proposed method was subjectively compared with the angle histogram and R-histogram and was evaluated as better with respect to results ranking. Zhiyun Xue, L. Rodney Long, Sameer K. Antani, George R. Thoma, Jose Jeronimo |
ICPR | 3 |
| 2008 | Automatic medical image annotation and retrieval
Jian Yao 0003, Zhongfei Zhang, Sameer K. Antani, L. Rodney Long, George R. Thoma |
Neurocomputing | 3 |
| 2008 | A Spine X-Ray Image Retrieval System Using Partial Shape MatchingabstractIn recent years, there has been a rapid increase in the size and number of medical image collections. Thus, the development of appropriate methods for medical information retrieval is especially important. In a large collection of spine X-ray images, maintained by the National Library of Medicine, vertebral boundary shape has been determined to be relevant to pathology of interest. This paper presents an innovative partial shape matching (PSM) technique using dynamic programming (DP) for the retrieval of spine X-ray images. The improved version of this technique called corner-guided DP is introduced. It uses nine landmark boundary points for DP search and improves matching speed by approximately 10 times compared to traditional DP. The retrieval accuracy and processing speed of the retrieval system based on the new corner-guided PSM method are evaluated and included in this paper. Xiaoqian Xu, Dah-Jye Lee, Sameer K. Antani, L. Rodney Long |
IEEE Trans. Inf. Technol. Biomed. | 3 |
| 2007 | Investigating CBIR Techniques for Cervicographic Images
Zhiyun Xue, Sameer K. Antani, L. Rodney Long, Jose Jeronimo, George R. Thoma |
AMIA | 2 |
| 2006 | Technology for Medical Education, Research, and Disease Screening by Exploitation of Biomarkers in a Large Collection of Uterine Cervix ImagesabstractThe Communications Engineering Branch of the National Library of Medicine is collaborating with the National Cancer Institute (NCI) in developing applications for medical education, research, and disease screening for precancer detection in the uterine cervix. These applications include (1) expert marking/labeling of tissue regions, (2) Web viewing/ interpretation of histology images, (3) image database/retrieval, and (4) training/testing in clinical image interpretation. Initial NCI studies have been conducted in expert cervicography marking and histology evaluation. We are working toward making cervix images searchable by content-based image retrieval (CBIR). Image pre-processing to remove specular reflection artifacts has achieved 90% success (120 images). Similar results have been obtained for automated location of cervix regions, using Gaussian Mixture Modeling (GMM) with Lab color and one geometric feature. We describe initial classification experiments to discriminate clinically significant tissue, using RGB, HSV, Lab, and YCbCr color models, texture measures, and GMM, Fuzzy C-means, and deterministic annealing algorithms. L. Rodney Long, Sameer K. Antani, Jose Jeronimo, Mark Schiffman, Mike Bopf, Leif Neve, Carl Cornwell, Scott R. Budihas, George R. Thoma |
CBMS | 2 |
| 2006 | Pre-Indexing for Fast Partial Shape Matching of Vertebrae ImagesabstractFast retrieval of images from large databases especially from large medical image databases is of great interest to researchers. We have developed shape-based image retrieval system for spine X-ray images that applies Dynamic Programming (DP) in Partial Shape Matching (PSM) techniques to vertebral boundary data. As we enable Internet access to this system with an aim to enable CBIR access to the entire NHANES II spine X-ray image collection, retrieval efficiency has become critical. In this paper, we present enhancements to our existing sequential retrieval model to provide faster partialshape- based vertebrae retrieval. Based on the characteristics of vertebral osteophyte pathology, anterior superior and anterior inferior parts are the areas with the most interest to the users and are indexed for each shape. Pair-wise distances between indexed parts are pre-calculated and used by agglomerative clustering algorithm to index the whole database. To increase the retrieval speed, PSM with DP is therefore conducted only on a small selected set of shapes by the pre-indexing retrieval. Xiaoqian Xu, Dah-Jye Lee, Sameer K. Antani, L. Rodney Long |
CBMS | 3 |
| 2006 | Automatic Medical Image Annotation and Retrieval Using SECCabstractThe demand for automatically annotating and retrieving medical images is growing faster than ever. In this paper, we present a novel medical image annotation method based on the proposed Semantic Error-Correcting output Codes (SECC). With this annotation method, we present a new semantic image retrieval method, which exploits the high level semantic similarity. For example, a user may query the system using an image of arm while he/she expects images of hand. This cannot be realized by traditional retrieval methods. The experimental results on the IMAGECLEF 2005 annotation data set clearly show the strength and the promise of the presented methods. Jian Yao 0003, Sameer K. Antani, L. Rodney Long, George R. Thoma, Zhongfei Zhang |
CBMS | 2 |
| 2006 | Automatic Medical Image Annotation and Retrieval using SEMI-SECCabstractThe demand for automatically annotating and retrieving medical images is growing faster than ever. In this paper, we present a novel medical image retrieval method based on SEMI-supervised Semantic Error-Correcting output Codes (SEMI-SEC). The experimental results on IMAGECLEF 2005 [1] annotation data set clearly show the strength and the promise of the presented methods. Jian Yao 0003, Zhongfei Zhang, Sameer K. Antani, L. Rodney Long, George R. Thoma |
ICME | 3 |
| 2005 | Relevance Feedback for Spine X-ray RetrievalabstractRelevance feedback (RF) has been an active research area in content-based image retrieval (CBIR). RF intends to bridge the gap between the low-level image features and the high-level human visual perception by analyzing and employing the feedback information provided by the user. This gap becomes more evident and important in medical image retrieval due to the two distinct facts with regard to medical images: (1) subtle differences between images, even between pathological and non-pathological images; (2) subjective and different diagnosis even among experts. This paper describes a novel linear weight-updating approach for RF applying to spine X-ray image retrieval. The algorithm utilizes both positive and negative examples to gain feedback from the user. Experimental results show that the proposed approach can substantially improve the retrieval performance to better satisfy the individual user's preferences. Xiaoqian Xu, Dah-Jye Lee, Sameer K. Antani, L. Rodney Long |
CBMS | 3 |
| 2004 | An Architecture for Streamlining the Implementation of Biomedical Text/Image Databases on the WebabstractTo date the implementation of biomedical text/images databases on the Web has been hindered by the lack of software flexibility to enable new datasets to be "rolled in " to existing database systems without extensive programming modifications. An R&D division of the National Library of Medicine (NLM) is creating a new database system, the multimedia database tool (MDT), built on an architecture that moves the required customization for a particular dataset from the programmer to the database administrator level, thereby reducing potentially high labor costs. The first database application to be supported will be uterine cervix data from the National Cancer Institute, including 100,000 digitized images, and NHANES II and III databases now hosted by the NLM Web-based medical information retrieval system (WebMIRS). Mike Bopf, Tina Coleman, L. Rodney Long, Sameer K. Antani, George R. Thoma, Jose Jeronimo, Mark Schiffman |
CBMS | 4 |
| 2004 | A Tool for Collection of Region Based Data from Uterine Cervix Images for Correlation of Visual and Clinical Variables Related to Cervical NeoplasiaabstractThe National Cancer Institute (NCI) is collaborating with the National Library of Medicine (NLM) to create a database of digitized images of the uterine cervix for research, training, and education. The database of 100,000 images collected in NCI projects will be Web-accessible and will contain not only the digitized images, but also clinical, longitudinal data relevant to the detection of uterine cervix cancer. The database will also contain spatial data in the form of expert-marked boundary regions of anatomy and tissue regions of high interest. The collection of this spatial data is enabled by customized software, the Boundary Marking Tool (BMT). The BMT, its significance, and its use in preliminary research studies, are described in this paper. Jose Jeronimo, Mark Schiffman, L. Rodney Long, Leif Neve, Sameer K. Antani |
CBMS | 5 |
| 2004 | Image Analysis Techniques for the Automated Evaluation of Subaxial Subluxation in Cervical Spine X-ray ImagesabstractRheumatoid arthritis is a chronic inflammatory disease affecting synovial joints of the body, especially the hands and feet, spine, knees and hips. For many patients, the cervical spine is associated with rheumatoid arthritis. Subluxation is the abnormal movement of one of the bones that comprise a joint. In this research, image analysis techniques have been investigated for the recognition of cervical spine x-ray images with one or more instances of subaxial subluxation. Receiver operating characteristic curve results are presented, showing potential for subaxial subluxation discrimination on an image-by-image basis. Ronald Joe Stanley, Santhosh Seetharaman, L. Rodney Long, Sameer K. Antani, George R. Thoma, Edward Downey |
CBMS | 4 |
| 2004 | Partial Shape Matching of Spine X-Ray Shapes Using Dynamic ProgrammingabstractThe osteophyte shows only on some particular locations on the vertebra. This indicates that other locations on the vertebra shape contain that are not of interest hinder the spine X-ray image retrieval relevance precision, which motivates our research in partial shape matching (PSM). This paper presents PSM methods for matching shapes with variable number of points and with different data point distributions. Dynamic programming (DP) is proposed for matching partial shapes by allowing merging of the data points in the process of PSM. DP is implemented based on two shape representation methods: line segments and multiple open triangles. The performance evaluation, which is based on human relevance judgements of these two shape representations in PSM is also presented. Xiaoqian Xu, Dah-Jye Lee, Sameer K. Antani, L. Rodney Long |
CBMS | 3 |
| 2004 | Shape based retrieval in NHANES IIabstractNHANES II is a nationally significant medical image database of spine x-ray images located at the National Library of Medicine. A key feature of spine disease in these images is the presence of osteoytes which are bony processes that alter the shape of vertebrae. Shapes of vertebrae are conveniently described in shape spaces which are non-linear manifolds. Indexing in such non-linear manifolds is an open problem. In this paper, we describe a technique of embedding shape manifolds in Euclidean spaces in a way that allows the use of classical indexing techniques for indexing shape. Application of this to the NHANES II database is also described. Hemant D. Tagare, Xiaoning Qian, Robert K. Fulbright, L. Rodney Long, Sameer K. Antani |
ACM Multimedia | 5 |
| 2004 | Evaluation of shape similarity measurement methods for spine X-ray images
Sameer K. Antani, Dah-Jye Lee, L. Rodney Long, George R. Thoma |
J. Vis. Commun. Image Represent. | 1 |
| 2003 | An Imaging System Correlating Lip Shapes with Tongue Contact Patterns for Speech Pathology ResearchabstractIn this research, an imaging system was built to work with a newly developed electronic device to help people produce sounds correctly. The system consists of two parts, the internal tongue contact pattern data collection and the external lip shape information analysis. The tongue position information was gathered using the palatometer, an innovative tongue contact pattern-tracking device invented by Dr. Samuel Fletcher. The lip shape information was collected by processing images taken from people articulating different sounds. We developed an efficient color image segmentation technique to extract lip contour points and form a closed curve for shape analysis. The geometry invariant turn function vs. normalized length was then calculated for the lip shape for each sound and compared against the turn function of the lips in a resting position to quantify their variations from this reference. Both internal (vocal tract) and external (visible lip shape) information was collected for each of the speech sounds. The lip shape information extracted from the images was then correlated with tongue position information. The test results showed that this imaging system can be used to quantify the lip shape information and its relations with the tongue position and is a potentially useful tool for speech pathology research. Dah-Jye Lee, Daniel Bates, Christopher Dromey, Xiaoqian Xu, Sameer K. Antani |
CBMS | 5 |
| 2003 | A Prototype Content-Based Image Retrieval System for Spine X-RaysabstractAt the Lister Hill National Center for Biomedical Communications, an R&D division of the National Library of Medicine, we are engaged in an effort in content-based image retrieval (CBIR) for biomedical image databases. Toward the goal of developing a functional and significant CBIR capability, we have created a prototype system for image indexing and retrieval which operates on a collection of spine X-rays and associated health survey data. In this paper, we present our prototype system functionality, performance results, ongoing research, and outstanding technical issues. L. Rodney Long, Sameer K. Antani, George R. Thoma |
CBMS | 2 |
| 2003 | Localizing Contour Points for Indexing an X-Ray Image Retrieval SystemabstractVertebra shape can effectively describe various pathologies found in spine X-ray images. There are some critical regions on the shape contour which help determine whether the shape is pathologic or normal. We selected a subset of 250 segmented vertebra boundaries for study from a collection of 17,000 digitized X-rays of cervical and lumbar spine taken as a part of the second National Health and Nutrition Examination Survey (NHANES II). A board certified expert radiologist marked nine morphometric landmark points on the contour of these cervical and lumbar images. Image indexing could mimic the model used by the radiologists to mark the images, e.g. 6-, 9-, or 10-point, thereby improve the query and retrieval of vertebra shapes from the image database. In this paper, we present a technique to automatically select nine points from the boundary contour. The comparison between two 9-point models using the L/sub 2/ distance and retrieval rank results derived respectively from the 9-point model marked by the expert and the 9-point model selected with our algorithm provides a good measure of how well the two models match. Xiaoqian Xu, Dah-Jye Lee, Sameer K. Antani, L. Rodney Long |
CBMS | 3 |
| 2003 | Vertebra shape classification using MLP for content-based image retrievalabstractA desirable content-based image retrieval (CBIR) system would classify extracted image features to support some form of semantic retrieval. The Lister Hill National Center for Biomedical Communications, an intramural R&D division of the National Library for Medicine (NLM), maintains an archive of digitized X-rays of the cervical and lumbar spine taken as part of the second national health and nutrition examination survey (NHANES II). It is our goal to provide shape-based access to digitized X-rays including retrieval on automatically detected and classified pathology, e.g., anterior osteophytes. This is done using radius of curvature analysis along the anterior portion, and morphological analysis for quantifying protrusion regions along the vertebra boundary. Experimental results are presented for the classification of 704 cervical spine vertebrae by evaluating the features using a multi-layer perceptron (MLP) based approach. In this paper, we describe the design and current status of the content-based image retrieval (CBIR) system and the role of neural networks in the design of an effective multimedia information retrieval system. Sameer K. Antani, L. Rodney Long, George R. Thoma, Ronald Joe Stanley |
IJCNN | 1 |
| 2003 | Extraction of special effects caption text events from digital video
David Crandall, Sameer K. Antani, Rangachar Kasturi |
Int. J. Document Anal. Recognit. | 2 |
| 2002 | A survey on the use of pattern recognition methods for abstraction, indexing and retrieval of images and video
Sameer K. Antani, Rangachar Kasturi, Ramesh Jain 0001 |
Pattern Recognit. | 1 |
| 2000 | Robust Extraction of Text in VideoabstractDespite advances in the archiving of digital video, we are still unable to efficiently search and retrieve the portions that interest us. Video indexing by shot segmentation has been a proposed solution and several research efforts are seen in the literature. Shot segmentation alone cannot solve the problem of content based access to video. Recognition of text in video has been proposed as an additional feature. Several research efforts are found in the literature for text extraction from complex images and video with applications for video indexing. We present an update of our system for detection and extraction of an unconstrained variety of text from general purpose video. The text detection results from a variety of methods are fused and each single text instance is segmented to enable it for OCR. Problems in segmenting text from video are similar to those faced in detection and localization phases. Video has low resolution and the text often has poor contrast with a changing background. The proposed system applies a variety of methods and takes advantage of the temporal redundancy in video resulting in good text segmentation. Sameer K. Antani, David Crandall, Rangachar Kasturi |
ICPR | 1 |
| 2000 | Application of Planar Motion Segmentation for Scene Text ExtractionabstractThis paper explores an approach for extracting scene text from a sequence of images with relative motion between the camera and the scene. It is assumed that the scene text lies on planar surfaces, whereas the other features are likely to be at random depths or undergoing independent motion. The motion model parameters of these planar surfaces are estimated using gradient based methods, and multiple motion segmentation. The equations of the planar surfaces, as well as the camera motion parameters are extracted by combining the motion models of multiple planar surfaces. This approach is expected to improve the reliability and robustness of the estimates, which are used to perform perspective correction on the individual surfaces. Perspective correction can lead to improvement in OCR performance. This work can be useful for detecting road signs and bill-boards from a moving vehicle. Tarak Gandhi, Rangachar Kasturi, Sameer K. Antani |
ICPR | 3 |
| 1999 | Gujarati Character RecognitionabstractThis paper describes the classification of a subset of printed or digitized Gujarati characters. Gujarati belongs to the genre of Devanagri scripts from the Indian subcontinent. Very little work is found in the literature for recognition of Indian language scripts. For this paper a subset of similar appearing Gujarati characters was chosen and subjected to classification by different classifiers. The sample and test images for the characters were obtained from digital images available on the Internet and from scanned images of printed Gujarati text. For their classification, the Euclidean Minimum Distance and the k-Nearest Neighbor classifiers were used with regular and invariant moments. The characters were also classified in the binary feature space using Hamming Distance classifier. The paper presents the recognition rates for these classifiers. A recognition rate of 67% is achieved. Sameer K. Antani, Lalitha Agnihotri |
ICDAR | 1 |
| 1999 | A System for Automatic Text Detection in VideoabstractVideo indexing is an important problem that has occupied recent research efforts. The text appearing in video can provide semantic information about the scene content. Detecting and recognizing text events can provide indices into the video for content based querying. We describe a system for detecting, tracking, and extracting artificial and scene text in MPEG-1 video. Preliminary results are presented. Ullas Gargi, David Crandall, Sameer K. Antani, Tarak Gandhi, Ryan Keener, Rangachar Kasturi |
ICDAR | 3 |
| 1998 | VADIS: A Video Analysis, Display and Indexing System
Ullas Gargi, Sameer K. Antani, Rangachar Kasturi |
CVPR | 2 |
| 1998 | Performance Characterization and Comparison of Video Indexing AlgorithmsabstractTemporal segmentation of video is a necessary first step to indexing digital video for browsing and retrieval. A number of different video temporal segmentation algorithms have been published in the literature. There has been little effort to evaluate and characterize their performance so as to deliver a single (or set of) algorithms that may be used by other researchers for indexing video databases. The present results of evaluating a number of these algorithms and characterizing their performance, specifically with respect to robustness to encoder and bitrate changes. The lessons learnt have relevance to algorithm development and evaluation in general. Ullas Gargi, Rangachar Kasturi, Sameer K. Antani |
CVPR | 3 |
| 1998 | Indexing text events in digital video databasesabstractLike shot changes, the presence of text in digital video is an important event that can be used to index digital video and provide extremely useful semantic information about the scene content. The special characteristics of digital video compared to document images both require and allow new robust approaches to recognition of text in video. We discuss the characteristics and special challenges of text in video and present a strategy of detecting, localizing, and segmenting text from video data for the text indexing problem. Preliminary results from our approach are presented. Ullas Gargi, Sameer K. Antani, Rangachar Kasturi |
ICPR | 2 |