EDBT 2026 Demo / reviewers in the wild / expert
Hagit Shatkay
dblp:63/2826
· DBLP profile ↗
50ranked-venue papers
10as first author
6since 2021 · last 2025
0000-0001-6953-3970ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 40 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BI-LAVA: Biocuration With Hierarchical Image Labelling Through Active Learning and Visual AnalyticsabstractAbstract In the biomedical domain, taxonomies organize the acquisition modalities of scientific images in hierarchical structures. Such taxonomies leverage large sets of correct image labels and provide essential information about the importance of a scientific publication, which could then be used in biocuration tasks. However, the hierarchical nature of the labels, the overhead of processing images, the absence or incompleteness of labelled data and the expertise required to label this type of data impede the creation of useful datasets for biocuration. From a multi‐year collaboration with biocurators and text‐mining researchers, we derive an iterative visual analytics and active learning (AL) strategy to address these challenges. We implement this strategy in a system called BI‐LAVA—Biocuration with Hierarchical Image Labelling through Active Learning and Visual Analytics. BI‐LAVA leverages a small set of image labels, a hierarchical set of image classifiers and AL to help model builders deal with incomplete ground‐truth labels, target a hierarchical taxonomy of image modalities and classify a large pool of unlabelled images. BI‐LAVA's front end uses custom encodings to represent data distributions, taxonomies, image projections and neighbourhoods of image thumbnails, which help model builders explore an unfamiliar image dataset and taxonomy and correct and generate labels. An evaluation with machine learning practitioners shows that our mixed human–machine approach successfully supports domain experts in understanding the characteristics of classes within the taxonomy, as well as validating and improving data quality in labelled and unlabelled collections. Juan Trelles Trabucco, Andrew Wentzel, William Berrios, Hagit Shatkay, G. Elisabeta Marai |
Comput. Graph. Forum | 4 |
| 2023 | MouseScholar: Evaluating an Image+Text Search System for BiocurationabstractBiocuration is the process of analyzing biological or biomedical articles to organize biological data into data repositories using taxonomies and ontologies. Due to the expanding number of articles and the relatively small number of biocurators, automation is desired to improve the workflow of assessing articles worth curating. As figures convey essential information, automatically integrating images may improve curation. In this work, we instantiate and evaluate a first-in-kind, hybrid image+text document search system for biocuration. The system, MouseScholar, leverages an image modality taxonomy derived in collaboration with biocurators, in addition to figure segmentation, and classifiers components as a back-end and a streamlined front-end interface to search and present document results. We formally evaluated the system with ten biocurators on a mouse genome informatics biocuration dataset and collected feedback. The results demonstrate the benefits of blending text and image information when presenting scientific articles for biocuration. Juan Trelles Trabucco, Carla Floricel, Cecilia N. Arighi, Hagit Shatkay, Daniela Raciti, Martin Ringwald, G. Elisabeta Marai |
BIBM | 4 |
| 2022 | Toward ECG-based analysis of hypertrophic cardiomyopathy: a novel ECG segmentation method for handling abnormalitiesabstractOBJECTIVE: Abnormalities in impulse propagation and cardiac repolarization are frequent in hypertrophic cardiomyopathy (HCM), leading to abnormalities in 12-lead electrocardiograms (ECGs). Computational ECG analysis can identify electrophysiological and structural remodeling and predict arrhythmias. This requires accurate ECG segmentation. It is unknown whether current segmentation methods developed using datasets containing annotations for mostly normal heartbeats perform well in HCM. Here, we present a segmentation method to effectively identify ECG waves across 12-lead HCM ECGs. METHODS: We develop (1) a web-based tool that permits manual annotations of P, P', QRS, R', S', T, T', U, J, epsilon waves, QRS complex slurring, and atrial fibrillation by 3 experts and (2) an easy-to-implement segmentation method that effectively identifies ECG waves in normal and abnormal heartbeats. Our method was tested on 131 12-lead HCM ECGs and 2 public ECG sets to evaluate its performance in non-HCM ECGs. RESULTS: Over the HCM dataset, our method obtained a sensitivity of 99.2% and 98.1% and a positive predictive value of 92% and 95.3% when detecting QRS complex and T-offset, respectively, significantly outperforming a state-of-the-art segmentation method previously employed for HCM analysis. Over public ECG sets, it significantly outperformed 3 state-of-the-art methods when detecting P-onset and peak, T-offset, and QRS-onset and peak regarding the positive predictive value and segmentation error. It performed at a level similar to other methods in other tasks. CONCLUSION: Our method accurately identified ECG waves in the HCM dataset, outperforming a state-of-the-art method, and demonstrated similar good performance as other methods in normal/non-HCM ECG sets. Kasra Nezamabadi, Jacob Mayfield, Pengyuan Li 0001, Gabriela V. Greenland, Sebastian Rodriguez 0001, Bahadir Simsek, Parvin Mousavi, Hagit Shatkay, M. Roselle Abraham |
J. Am. Medical Informatics Assoc. | 8 |
| 2021 | ANIMO: Annotation of Biomed Image ModalitiesabstractFigures within biomedical articles present essential evidence of the relevance of a publication in a curation workflow. In particular, visual cues of the image modality or experimental methods can help expert curators identify relevant papers from an increasing number of publications. Automating the identification of these content-bearing images can thus be helpful in computer-assisted curation. However, the paucity of labeled datasets and the specialized training required to label such images hinder the development of such tools. To address this problem, we present the design of ANIMO, a labeling system that integrates extraction and segmentation tools to ease the annotation burden. We first introduce two taxonomies of image modalities and experimental methods, derived in collaboration with curators. On the back-end of the system, we process batches of documents and create a labeling task per document. At the front-end, expert curators can access these tasks through a web interface and access the article of interest. We describe the evaluation of this system by a group of biocurators, and the human factor lessons learned from this interdisciplinary experience. Juan Trelles Trabucco, Pengyuan Li 0001, Cecilia N. Arighi, Daniela Raciti, Hagit Shatkay, G. Elisabeta Marai |
BIBM | 5 |
| 2021 | Corrigendum to: Utilizing image and caption information for biomedical document classificationabstractBioinformatics (2021), Volume 37(Suppl1), i468–i476, doi:10.1093/bioinformatics/btab331 The error is thus only in mis-typing the formulae themselves; not in the actual calculations of the precision and the recall used throughout the paper. As such, the other parts of the manuscript – specifically the experimental results reported, all stand as they appear in the original publication, and are not impacted by this correction. Pengyuan Li 0001, Xiangying Jiang, Juan Trelles Trabucco, Daniela Raciti, Cynthia L. Smith, Martin Ringwald, G. Elisabeta Marai, Cecilia N. Arighi, Hagit Shatkay |
Bioinform. | 10 |
| 2021 | Utilizing image and caption information for biomedical document classificationabstractMOTIVATION: Biomedical research findings are typically disseminated through publications. To simplify access to domain-specific knowledge while supporting the research community, several biomedical databases devote significant effort to manual curation of the literature-a labor intensive process. The first step toward biocuration requires identifying articles relevant to the specific area on which the database focuses. Thus, automatically identifying publications relevant to a specific topic within a large volume of publications is an important task toward expediting the biocuration process and, in turn, biomedical research. Current methods focus on textual contents, typically extracted from the title-and-abstract. Notably, images and captions are often used in publications to convey pivotal evidence about processes, experiments and results. RESULTS: We present a new document classification scheme, using both image and caption information, in addition to titles-and-abstracts. To use the image information, we introduce a new image representation, namely Figure-word, based on class labels of subfigures. We use word embeddings for representing captions and titles-and-abstracts. To utilize all three types of information, we introduce two information integration methods. The first combines Figure-words and textual features obtained from captions and titles-and-abstracts into a single larger vector for document representation; the second employs a meta-classification scheme. Our experiments and results demonstrate the usefulness of the newly proposed Figure-words for representing images. Moreover, the results showcase the value of Figure-words, captions and titles-and-abstracts in providing complementary information for document classification; these three sources of information when combined, lead to an overall improved classification performance. AVAILABILITY AND IMPLEMENTATION: Source code and the list of PMIDs of the publications in our datasets are available upon request. Pengyuan Li 0001, Xiangying Jiang, Juan Trelles Trabucco, Daniela Raciti, Cynthia L. Smith, Martin Ringwald, G. Elisabeta Marai, Cecilia N. Arighi, Hagit Shatkay |
Bioinform. | 10 |
| 2020 | Modality-Classification of Microscopy Images Using Shallow Variants of Deep NetworksabstractMicroscopy images are pervasive in biomedical research publications, where images obtained through various microscopy modalities (light, fluorescence, scanning, transmission) are often used to describe and summarize experiments and contributions. Hence, there is growing interest in automatically identifying these microscopy images' modality and utilizing this knowledge in automated search tools. However, identifying microscopy images poses challenges due to a lack of extensive collections of labeled images. We describe and evaluate two alternative approaches to microscopy image classification. In the first approach, we progressively fine-tuned layers of ResNet models. The second approach uses shallow variants of ResNet networks, where we leverage the outputs from previous convolutional blocks. We compare these results against a Support Vector Machine (SVM)-based baseline. Our results show that fine-tuning specific layers yields better results than fine-tuning the whole model. Furthermore, shallower variants produce competitive results when compared to the entire fine-tuned model. Juan Trelles Trabucco, Pengyuan Li 0001, Cecilia N. Arighi, Hagit Shatkay, G. Elisabeta Marai |
BIBM | 4 |
| 2019 | Figure and caption extraction from biomedical documentsabstractMOTIVATION: Figures and captions convey essential information in biomedical documents. As such, there is a growing interest in mining published biomedical figures and in utilizing their respective captions as a source of knowledge. Notably, an essential step underlying such mining is the extraction of figures and captions from publications. While several PDF parsing tools that extract information from such documents are publicly available, they attempt to identify images by analyzing the PDF encoding and structure and the complex graphical objects embedded within. As such, they often incorrectly identify figures and captions in scientific publications, whose structure is often non-trivial. The extraction of figures, captions and figure-caption pairs from biomedical publications is thus neither well-studied nor yet well-addressed. RESULTS: We introduce a new and effective system for figure and caption extraction, PDFigCapX. Unlike existing methods, we first separate between text and graphical contents, and then utilize layout information to effectively detect and extract figures and captions. We generate files containing the figures and their associated captions and provide those as output to the end-user.We test our system both over a public dataset of computer science documents previously used by others, and over two newly collected sets of publications focusing on the biomedical domain. Our experiments and results comparing PDFigCapX to other state-of-the-art systems show a significant improvement in performance, and demonstrate the effectiveness and robustness of our approach. AVAILABILITY AND IMPLEMENTATION: Our system is publicly available for use at: https://www.eecis.udel.edu/~compbio/PDFigCapX. The two new datasets are available at: https://www.eecis.udel.edu/~compbio/PDFigCapX/Downloads. Pengyuan Li 0001, Xiangying Jiang, Hagit Shatkay |
Bioinform. | 3 |
| 2018 | Extracting Figures and Captions from Scientific PublicationsabstractFigures and captions convey essential information in scientific publications. As such, there is a growing interest in mining published figures and in utilizing their respective captions as a source of knowledge. There is also much interest in image captioning systems that can automatically generate captions for images, whose training requires large datasets of image-caption pairs. Notably, the first fundamental step of obtaining figures and captions from publications is neither well-studied nor yet well-addressed. In this paper, we introduce a new and effective system for figure and caption extraction, PDFigCapX. Unlike current methods that extract figures by handling raw encoded contents of PDF documents, we separate text from graphical contents and utilize layout information to detect and disambiguate figures and captions. Files containing the figures and their associated captions are then produced as output to the end-user. We test PDFigCapX on both a previously used generic dataset and on two new sets of publications within the biomedical domain. Our experiments and results show a significant improvement in performance compared to the state-of-the-art, and demonstrate the effectiveness of our approach. Our system will be available for use at: https://www.eecis.udel.edu/~compbio/PDFigCapX. Pengyuan Li 0001, Xiangying Jiang, Hagit Shatkay |
CIKM | 3 |
| 2018 | Compound image segmentation of published biomedical figuresabstractMotivation: Images convey essential information in biomedical publications. As such, there is a growing interest within the bio-curation and the bio-databases communities, to store images within publications as evidence for biomedical processes and for experimental results. However, many of the images in biomedical publications are compound images consisting of multiple panels, where each individual panel potentially conveys a different type of information. Segmenting such images into constituent panels is an essential first step toward utilizing images. Results: In this article, we develop a new compound image segmentation system, FigSplit, which is based on Connected Component Analysis. To overcome shortcomings typically manifested by existing methods, we develop a quality assessment step for evaluating and modifying segmentations. Two methods are proposed to re-segment the images if the initial segmentation is inaccurate. Experimental results show the effectiveness of our method compared with other methods. Availability and implementation: The system is publicly available for use at: https://www.eecis.udel.edu/~compbio/FigSplit. The code is available upon request. Contact: [email protected]. Supplementary information: Supplementary data are available online at Bioinformatics. Pengyuan Li 0001, Xiangying Jiang, Chandra Kambhamettu, Hagit Shatkay |
Bioinform. | 4 |
| 2018 | Co-occurrence of medical conditions: Exposing patterns through probabilistic topic modeling of snomed codesabstractPatients associated with multiple co-occurring health conditions often face aggravated complications and less favorable outcomes. Co-occurring conditions are especially prevalent among individuals suffering from kidney disease, an increasingly widespread condition affecting 13% of the general population in the US. This study aims to identify and characterize patterns of co-occurring medical conditions in patients employing a probabilistic framework. Specifically, we apply topic modeling in a non-traditional way to find associations across SNOMED-CT codes assigned and recorded in the EHRs of >13,000 patients diagnosed with kidney disease. Unlike most prior work on topic modeling, we apply the method to codes rather than to natural language. Moreover, we quantitatively evaluate the topics, assessing their tightness and distinctiveness, and also assess the medical validity of our results. Our experiments show that each topic is succinctly characterized by a few highly probable and unique disease codes, indicating that the topics are tight. Furthermore, inter-topic distance between each pair of topics is typically high, illustrating distinctiveness. Last, most coded conditions grouped together within a topic, are indeed reported to co-occur in the medical literature. Notably, our results uncover a few indirect associations among conditions that have hitherto not been reported as correlated in the medical literature. Moumita Bhattacharya, Claudine Jurkovitz, Hagit Shatkay |
J. Biomed. Informatics | 3 |
| 2017 | Assessing chronic kidney disease from office visit records using hierarchical meta-classification of an imbalanced datasetabstractChronic Kidney Disease (CKD) is an increasingly prevalent condition affecting 13% of the US population. The disease is often a silent condition, making its diagnosis challenging. Identifying CKD stages from standard office visit records can help in early detection of the disease and lead to timely intervention. The dataset we use is highly imbalanced. We propose a hierarchical meta-classification method, aiming to stratify CKD by severity levels, employing simple quantitative non-text features gathered from office visit records, while addressing data imbalance. Our method effectively stratifies CKD severity levels obtaining high average sensitivity, precision and F-measure (~93%). We also conduct experiments in which the dimensionality of the data is significantly reduced to include only the most salient features. Our results show that the good performance of our system is retained even when using the reduced feature sets, as well as under much reduced training sets, indicating that our method is stable and generalizable. Moumita Bhattacharya, Claudine Jurkovitz, Hagit Shatkay |
BIBM | 3 |
| 2017 | Identifying articles relevant to drug-drug interaction: Addressing class imbalanceabstractInteractions between drugs (also known as drug-drug interactions or DDIs), which may cause adverse affects, are of much concern; predicting, anticipating and avoiding them is key for improving patient safety and treatment outcome. Knowledge of DDIs is important for physicians to avoid adverse effects when prescribing two drugs simultaneously. DDIs are often published in the biomedical literature; however, gathering information about DDIs is time consuming given the shear volume of publications. Automatic text classification can speed up access to documents related to DDIs. However, the biomedical literature contains a relatively small number of publications relevant to DDIs, compared to the vast amount of irrelevant publications. This imbalance can lead to incorrect classification. While methods addressing class imbalance have been introduced to correctly identify items in the minority (relevant) class to improve recall, they often misclassify items in the majority (irrelevant) class, which leads to low precision. To reduce the number of irrelevant documents misclassified as relevant (false positive), we develop a two-stage cascade classifier. In each step, we separate publication abstracts that are DDI-relevant from those that are either drug-irrelevant or drug-relevant but DDI-irrelevant. We compare our classifier with other popular learning methods that aim to handle imbalance, applying the methods to a well-curated corpus consisting of DDI-relevant and DDI-irrelevant PubMed abstracts. Our method achieves higher precision and F1 measure than other methods while maintaining similar recall. Moumita Bhattacharya, Heng-Yi Wu, Pengyuan Li 0001, Lang Li 0001, Hagit Shatkay |
BIBM | 6 |
| 2016 | Identifying patterns of associated-conditions through topic models of Electronic Medical RecordsabstractMultiple adverse health conditions co-occurring in a patient are typically associated with poor prognosis and increased office or hospital visits. Developing methods to identify patterns of co-occurring conditions can assist in diagnosis. Thus, identifying patterns of association among co-occurring conditions is of growing interest. In this paper, we report preliminary results from a data-driven study, in which we apply a machine learning method, namely, topic modeling, to Electronic Medical Records (EMRs), aiming to identify patterns of associated conditions. Specifically, we use the well-established Latent Dirichlet Allocation (LDA), a method based on the idea that documents can be modeled as a mixture of latent topics, where each topic is a distribution over words. In our study, we adapt the LDA model to identify latent topics in patients' EMRs. We evaluate the performance of our method both qualitatively and quantitatively, and show that the obtained topics indeed align well with distinct medical phenomena characterized by co-occurring conditions. Moumita Bhattacharya, Claudine Jurkovitz, Hagit Shatkay |
BIBM | 3 |
| 2016 | Improved Multi-Label Classification Using Inter-Dependence Structure via a Generative Mixture ModelabstractSingle-label classification associates each instance with a single label, while multi-label classification (MLC), assigns multiple labels to instances. Simple MLC systems assume that labels are independent of one another, while more complex approaches capture inter-dependencies among labels. Experiments comparing performance of MLC systems demonstrate that there is much room for improvement. Ramanuja Simha, Hagit Shatkay |
ECAI | 2 |
| 2016 | Prostate Cancer: Improved Tissue Characterization by Temporal Modeling of Radio-Frequency Ultrasound Echo Data
Layan Nahlawi, Farhad Imani, Mena Gaed, Jose A. Gomez, Madeleine Moussa, Eli Gibson, Aaron Fenster, Aaron D. Ward, Purang Abolmaesumi, Hagit Shatkay, Parvin Mousavi |
MICCAI (1) | 10 |
| 2015 | Using Hidden Markov Models to capture temporal aspects of ultrasound data in prostate cancerabstractRecent studies highlight temporal ultrasound data as highly promising in differentiating between malignant and benign tissues in prostate cancer patients. Since Hidden Markov Models can be used for capturing order and patterns in time varying signals, we employ them to model temporal aspects of ultrasound data that are typically not incorporated in existing models. By comparing order-preserving and order-altering models, we demonstrate that the order encoded in the series is necessary to model the variability in ultrasound data of prostate tissues. In future studies, we will investigate the influence of order on the differentiation between malignant and benign tissues. Layan Nahlawi, Farhad Imani, Mena Gaed, Jose A. Gomez, Madeleine Moussa, Eli Gibson, Aaron Fenster, Aaron D. Ward, Purang Abolmaesumi, Parvin Mousavi, Hagit Shatkay |
BIBM | 11 |
| 2015 | Utilizing image-based features in biomedical document classificationabstractImages form a rich information source, which remains underutilized in biomedical document classification. We present here work that uses both image- and text-based features in order to identify articles of interest, in this case, pertaining to cis-regulatory modules in the context of gene-networks. Extending on our new idea, which we have recently introduced, of using OCR-based features to identify DNA contents in images, we combine image and text based classifiers to categorize documents as relevant or irrelevant to cis-regulatory modules. Using a set of hundreds of articles, marked by experts as relevant or irrelevant to cis-regulatory modules, we train/test image and text based classifiers, as well as classifiers integrating both. Our results indicate that the latter show the best performance with Recall, F-measure and Utility measures all above 0.9, demonstrating the significance of incorporating image data, and specifically OCR-based features, into the document categorization process. Moreover, the use of character distribution properties to represent images is directly relevant to other biomedical images containing text (e.g. RNA, proteins). Diagrams and other images containing text are also prevalent outside the biomedical domain, hence the work stands to be applicable and beneficial in other application areas. Kaidi Ma, Hogyeong Jeong, M. V. Rohith, Gowri Somanath, Ryan Tarpine, Kyle Schutter, Dorothea Blostein, Sorin Istrail, Chandra Kambhamettu, Hagit Shatkay |
ICIP | 10 |
| 2015 | Protein (multi-)location prediction: utilizing interdependencies via a generative modelabstractMOTIVATION: Proteins are responsible for a multitude of vital tasks in all living organisms. Given that a protein's function and role are strongly related to its subcellular location, protein location prediction is an important research area. While proteins move from one location to another and can localize to multiple locations, most existing location prediction systems assign only a single location per protein. A few recent systems attempt to predict multiple locations for proteins, however, their performance leaves much room for improvement. Moreover, such systems do not capture dependencies among locations and usually consider locations as independent. We hypothesize that a multi-location predictor that captures location inter-dependencies can improve location predictions for proteins. RESULTS: We introduce a probabilistic generative model for protein localization, and develop a system based on it-which we call MDLoc-that utilizes inter-dependencies among locations to predict multiple locations for proteins. The model captures location inter-dependencies using Bayesian networks and represents dependency between features and locations using a mixture model. We use iterative processes for learning model parameters and for estimating protein locations. We evaluate our classifier MDLoc, on a dataset of single- and multi-localized proteins derived from the DBMLoc dataset, which is the most comprehensive protein multi-localization dataset currently available. Our results, obtained by using MDLoc, significantly improve upon results obtained by an initial simpler classifier, as well as on results reported by other top systems. AVAILABILITY AND IMPLEMENTATION: MDLoc is available at: http://www.eecis.udel.edu/∼compbio/mdloc. Ramanuja Simha, Sebastian Briesemeister, Oliver Kohlbacher, Hagit Shatkay |
Bioinform. | 4 |
| 2015 | Summary of the BioLINK SIG 2013 meeting at ISMB/ECCB 2013abstractUNLABELLED: The ISMB Special Interest Group on Linking Literature, Information and Knowledge for Biology (BioLINK) organized a one-day workshop at ISMB/ECCB 2013 in Berlin, Germany. The theme of the workshop was 'Roles for text mining in biomedical knowledge discovery and translational medicine'. This summary reviews the outcomes of the workshop. Meeting themes included concept annotation methods and applications, extraction of biological relationships and the use of text-mined data for biological data analysis. AVAILABILITY AND IMPLEMENTATION: All articles are available at http://biolinksig.org/proceedings-online/. Karin Verspoor, Hagit Shatkay, Lynette Hirschman, Christian Blaschke, Alfonso Valencia |
Bioinform. | 2 |
| 2014 | Identifying growth-patterns in children by applying cluster analysis to electronic medical recordsabstractObesity is one of the leading health concerns in the United States. Researchers and health care providers are interested in understanding factors affecting obesity and detecting the likelihood of obesity as early as possible. In this paper, we set out to recognize children who have higher risk of obesity by identifying distinct growth patterns in them. This is done by using clustering methods, which group together children who share similar body measurements over a period of time. The measurements characterizing children within the same cluster are plotted as a function of age. We refer to these plots as growth-pattern curves. We show that distinct growth-pattern curves are associated with different clusters and thus can be used to separate children into the topmost (heaviest), middle, or bottom-most cluster based on early growth measurements. Moumita Bhattacharya, Deborah Ehrenthal, Hagit Shatkay |
BIBM | 3 |
| 2014 | Using unsupervised learning to determine risk level for left ventricular diastolic dysfunctionabstractLeft Ventricular Diastolic Dysfunction (LVDD) is a decompensatory change in the relaxation properties of the heart, the risk for which increases with age. Currently, physicians use a decision-tree-like algorithm to distinguish between discrete LVDD levels. This approach, based on cut-off thresholds, can potentially lead to information loss and possibly to misdiagnosis. This paper aims to explore an alternative diagnostic method to determine LVDD risk level, taking into account a wide variety of attributes available in patient records, without pre-setting cut-off thresholds. Using a large dataset derived from the Baltimore Longitude Study of Aging (BLSA), and adjusting the data for age and gender, we employ the Chi Square test and the information gain criterion to identify attributes that correlate well with the physician-assigned grades; such attributes are referred to as distinguishing attributes. We then apply the expectation maximization (EM) algorithm, as well as the K-Means, in order to cluster records that are represented using distinguishing attributes. While clusters resulting from the K-Means are not stable, three stable and tightly-formed clusters, which are obtained from the EM algorithm, roughly correspond to the physician-assigned categories. Based on the results from the EM algorithm, we can compute a patient's probability to have low, high or no risk for LVDD, and use this probability as a basis for defining a risk score to determine the patient's LVDD severity. Kaidi Ma, Marco Canepa, James B. Strait, Hagit Shatkay |
BIBM | 4 |
| 2014 | Identifying hypertrophic cardiomyopathy patients by classifying individual heartbeats from 12-lead ECG signalsabstractTest based on electrocardiograms (ECG) that record the heart electrical activity can help in early detection of patients with hypertrophic cardiomyopathy (HCM) where the heart muscle is partially thickened and blood flow is (potentially fatally) obstructed. This paper presents a cardiovascular-patient classifier we developed to identify HCM patients using standard 10-seconds, 12-lead ECG signals. Patients are classified as having HCM if the majority of the heartbeats are recognized as HCM. Thus, the classifier's underlying task is to recognize individual heartbeats segmented from 12-lead ECG signals as HCM beats, where heartbeats from non-HCM cardiovascular patients are used as controls. We extracted 504 morphological and temporal features - both commonly used and newly-developed ones - from ECG signals for heartbeat classification. To assess classification performance, we trained and tested a random forest classifier and a support vector machine classifier using 5-fold cross validation. The patient-classification precision and F-measure of both classifiers are close to 0.85. Recall (sensitivity) and specificity are approximately 0.90. We also conducted feature selection experiments by gradually removing the least informative features; the results show that a relatively small subset of 304 highly informative features can achieve performance measures comparable to that achieved by using the complete set of features. Quazi Abidur Rahman, Larisa G. Tereshchenko, Matthew Kongkatong, Theodore Abraham, M. Roselle Abraham, Hagit Shatkay |
BIBM | 6 |
| 2013 | Protein (Multi-)Location Prediction: Using Location Inter-dependencies in a Probabilistic Framework
Ramanuja Simha, Hagit Shatkay |
WABI | 2 |
| 2013 | Protein Function Prediction using Text-based Features extracted from the Biomedical Literature: The CAFA ChallengeabstractBACKGROUND: Advances in sequencing technology over the past decade have resulted in an abundance of sequenced proteins whose function is yet unknown. As such, computational systems that can automatically predict and annotate protein function are in demand. Most computational systems use features derived from protein sequence or protein structure to predict function. In an earlier work, we demonstrated the utility of biomedical literature as a source of text features for predicting protein subcellular location. We have also shown that the combination of text-based and sequence-based prediction improves the performance of location predictors. Following up on this work, for the Critical Assessment of Function Annotations (CAFA) Challenge, we developed a text-based system that aims to predict molecular function and biological process (using Gene Ontology terms) for unannotated proteins. In this paper, we present the preliminary work and evaluation that we performed for our system, as part of the CAFA challenge. RESULTS: We have developed a preliminary system that represents proteins using text-based features and predicts protein function using a k-nearest neighbour classifier (Text-KNN). We selected text features for our classifier by extracting key terms from biomedical abstracts based on their statistical properties. The system was trained and tested using 5-fold cross-validation over a dataset of 36,536 proteins. System performance was measured using the standard measures of precision, recall, F-measure and overall accuracy. The performance of our system was compared to two baseline classifiers: one that assigns function based solely on the prior distribution of protein function (Base-Prior) and one that assigns function based on sequence similarity (Base-Seq). The overall prediction accuracy of Text-KNN, Base-Prior, and Base-Seq for molecular function classes are 62%, 43%, and 58% while the overall accuracy for biological process classes are 17%, 11%, and 28% respectively. Results obtained as part of the CAFA evaluation itself on the CAFA dataset are reported as well. CONCLUSIONS: Our evaluation shows that the text-based classifier consistently outperforms the baseline classifier that is based on prior distribution, and typically has comparable performance to the baseline classifier that uses sequence similarity. Moreover, the results suggest that combining text features with other types of features can potentially lead to improved prediction performance. The preliminary results also suggest that while our text-based classifier can be used to predict both molecular function and biological process in which a protein is involved, the classifier performs significantly better for predicting molecular function than for predicting biological process. A similar trend was observed for other classifiers participating in the CAFA challenge. Hagit Shatkay |
BMC Bioinform. | 2 |
| 2012 | Building a classifier for identifying sentences pertaining to disease-drug relationships in tardive dyskinesiaabstractIn this paper, we attempt to build a pipeline that identifies and extracts disease-drug relationships via sentence classification, and demonstrate the feasibility and utility of our approach using tardive dyskinesia as a case study. We manually developed and annotated a biomedicai training corpus for tardive dyskinesia. Using 10-fold cross validation, we tested and trained a naïve Bayes classifier to identify sentences pertaining to disease-drug relationships. Our precision, recall, and F-measure were all approximately 66%, and area under the ROC curve was over 80%. Our method helps to elucidate various drug effects on tardive dyskinesia and constitutes an initial effort toward the task of disease-drug relationship extraction. Xia Bi, Hongzhan Huang, Sherri Matis-Mitchell, Peter B. McGarvey, Manabu Torii, Hagit Shatkay, Cathy H. Wu |
BIBM | 6 |
| 2012 | Correction: A linear classifier based on entity recognition tools and a statistical approach to method extraction in the protein-protein interaction literatureabstractAbstract Correction to A. Lourenço, M. Conover, A. Wong, A. Nematzadeh, F. Pan, H. Shatkay, and L.M. Rocha."A Linear Classifier Based on Entity Recognition Tools and a Statistical Approach to Method Extraction in the Protein-Protein Interaction Literature". BMC Bioinformatics 2011, 12(Suppl 8):S12. doi: http://10.1186/1471-2105-12-S8-S12 . Anália Lourenço, Michael D. Conover, Azadeh Nematzadeh, Fengxia Pan, Hagit Shatkay, Luis M. Rocha |
BMC Bioinform. | 6 |
| 2011 | An Automatic System for Extracting Figures and Captions in Biomedical PDF DocumentsabstractFigures in biomedical articles often constitute direct evidence of experimental results. Image analysis methods can be coupled with text-based methods to improve knowledge discovery. However, automatically harvesting figures along with their associated captions from full-text articles remains challenging. In this paper, we present an automatic system for robustly harvesting figures from biomedical literature. Our approach relies on the idea that the PDF specification of the document layout can be used to identify encoded figures and figure boundaries within the PDF and enforce constraints among figure-regions. This allows us to harvest fragments of figures (subflgures), from the PDF, correctly identify subfigures that belong to the same figure, and identify the captions associated with each figure. Our method simultaneously recovers figures and captions and applies additional filtering process to remove irrelevant figures such as logos, to eliminate text passages that were incorrectly identified as captions, and to re-group subflgures to generate a putative figure. Finally, we associate figures with captions. Our preliminary experiments suggest that our method achieves an accuracy of 95% in harvesting figures-caption pairs from a set of 2,035 full-text biomedical documents from BioCreative III, containing 12,574 figures. Luis D. Lopez, Jingyi Yu 0001, Cecilia N. Arighi, Hongzhan Huang, Hagit Shatkay, Cathy H. Wu |
BIBM | 5 |
| 2011 | The Protein-Protein Interaction tasks of BioCreative III: classification/ranking of articles and linking bio-ontology concepts to full textabstractBACKGROUND: Determining usefulness of biomedical text mining systems requires realistic task definition and data selection criteria without artificial constraints, measuring performance aspects that go beyond traditional metrics. The BioCreative III Protein-Protein Interaction (PPI) tasks were motivated by such considerations, trying to address aspects including how the end user would oversee the generated output, for instance by providing ranked results, textual evidence for human interpretation or measuring time savings by using automated systems. Detecting articles describing complex biological events like PPIs was addressed in the Article Classification Task (ACT), where participants were asked to implement tools for detecting PPI-describing abstracts. Therefore the BCIII-ACT corpus was provided, which includes a training, development and test set of over 12,000 PPI relevant and non-relevant PubMed abstracts labeled manually by domain experts and recording also the human classification times. The Interaction Method Task (IMT) went beyond abstracts and required mining for associations between more than 3,500 full text articles and interaction detection method ontology concepts that had been applied to detect the PPIs reported in them. RESULTS: A total of 11 teams participated in at least one of the two PPI tasks (10 in ACT and 8 in the IMT) and a total of 62 persons were involved either as participants or in preparing data sets/evaluating these tasks. Per task, each team was allowed to submit five runs offline and another five online via the BioCreative Meta-Server. From the 52 runs submitted for the ACT, the highest Matthew's Correlation Coefficient (MCC) score measured was 0.55 at an accuracy of 89% and the best AUC iP/R was 68%. Most ACT teams explored machine learning methods, some of them also used lexical resources like MeSH terms, PSI-MI concepts or particular lists of verbs and nouns, some integrated NER approaches. For the IMT, a total of 42 runs were evaluated by comparing systems against manually generated annotations done by curators from the BioGRID and MINT databases. The highest AUC iP/R achieved by any run was 53%, the best MCC score 0.55. In case of competitive systems with an acceptable recall (above 35%) the macro-averaged precision ranged between 50% and 80%, with a maximum F-Score of 55%. CONCLUSIONS: The results of the ACT task of BioCreative III indicate that classification of large unbalanced article collections reflecting the real class imbalance is still challenging. Nevertheless, text-mining tools that report ranked lists of relevant articles for manual selection can potentially reduce the time needed to identify half of the relevant articles to less than 1/4 of the time when compared to unranked results. Detecting associations between full text articles and interaction detection method PSI-MI terms (IMT) is more difficult than might be anticipated. This is due to the variability of method term mentions, errors resulting from pre-processing of articles provided as PDF files, and the heterogeneity and different granularity of method term concepts encountered in the ontology. However, combining the sophisticated techniques developed by the participants with supporting evidence strings derived from the articles for human interpretation could result in practical modules for biological annotation workflows. Martin Krallinger, Miguel Vázquez, Florian Leitner, David Salgado, Andrew Chatr-aryamontri, Andrew G. Winter, Livia Perfetto, Leonardo Briganti, Luana Licata, Marta Iannuccelli, Luisa Castagnoli, Gianni Cesareni, Mike Tyers, Gerold Schneider, Fabio Rinaldi 0001, Robert Leaman, Graciela Gonzalez-Hernandez, Sérgio Matos, Sun Kim, W. John Wilbur, Luis M. Rocha, Hagit Shatkay, Ashish V. Tendulkar, Shashank Agarwal, Xinglong Wang, Rafal Rak, Keith Noto, Charles Elkan, Zhiyong Lu |
BMC Bioinform. | 22 |
| 2011 | A linear classifier based on entity recognition tools and a statistical approach to method extraction in the protein-protein interaction literatureabstractBACKGROUND: We participated, as Team 81, in the Article Classification and the Interaction Method subtasks (ACT and IMT, respectively) of the Protein-Protein Interaction task of the BioCreative III Challenge. For the ACT, we pursued an extensive testing of available Named Entity Recognition and dictionary tools, and used the most promising ones to extend our Variable Trigonometric Threshold linear classifier. Our main goal was to exploit the power of available named entity recognition and dictionary tools to aid in the classification of documents relevant to Protein-Protein Interaction (PPI). For the IMT, we focused on obtaining evidence in support of the interaction methods used, rather than on tagging the document with the method identifiers. We experimented with a primarily statistical approach, as opposed to employing a deeper natural language processing strategy. In a nutshell, we exploited classifiers, simple pattern matching for potential PPI methods within sentences, and ranking of candidate matches using statistical considerations. Finally, we also studied the benefits of integrating the method extraction approach that we have used for the IMT into the ACT pipeline. RESULTS: For the ACT, our linear article classifier leads to a ranking and classification performance significantly higher than all the reported submissions to the challenge in terms of Area Under the Interpolated Precision and Recall Curve, Mathew's Correlation Coefficient, and F-Score. We observe that the most useful Named Entity Recognition and Dictionary tools for classification of articles relevant to protein-protein interaction are: ABNER, NLPROT, OSCAR 3 and the PSI-MI ontology. For the IMT, our results are comparable to those of other systems, which took very different approaches. While the performance is not very high, we focus on providing evidence for potential interaction detection methods. A significant majority of the evidence sentences, as evaluated by independent annotators, are relevant to PPI detection methods. CONCLUSIONS: For the ACT, we show that the use of named entity recognition tools leads to a substantial improvement in the ranking and classification of articles relevant to protein-protein interaction. Thus, we show that our substantially expanded linear classifier is a very competitive classifier in this domain. Moreover, this classifier produces interpretable surfaces that can be understood as "rules" for human understanding of the classification. We also provide evidence supporting certain named entity recognition tools as beneficial for protein-interaction article classification, or demonstrating that some of the tools are not beneficial for the task. In terms of the IMT task, in contrast to other participants, our approach focused on identifying sentences that are likely to bear evidence for the application of a PPI detection method, rather than on classifying a document as relevant to a method. As BioCreative III did not perform an evaluation of the evidence provided by the system, we have conducted a separate assessment, where multiple independent annotators manually evaluated the evidence produced by one of our runs. Preliminary results from this experiment are reported here and suggest that the majority of the evaluators agree that our tool is indeed effective in detecting relevant evidence for PPI detection methods. Regarding the integration of both tasks, we note that the time required for running each pipeline is realistic within a curation effort, and that we can, without compromising the quality of the output, reduce the time necessary to extract entities from text for the ACT pipeline by pre-selecting candidate relevant text using the IMT pipeline. Anália Lourenço, Michael D. Conover, Azadeh Nematzadeh, Fengxia Pan, Hagit Shatkay, Luis M. Rocha |
BMC Bioinform. | 6 |
| 2010 | DTMBIO workshop summaryabstractNo abstract available. Hagit Shatkay, Doheon Lee, Min Song 0001, Shamkant B. Navathe |
CIKM | 1 |
| 2009 | An integrative scoring system for ranking SNPs by their potential deleterious effectsabstractMOTIVATION: Identifying single nucleotide polymorphisms (SNPs) that underlie common and complex human diseases, such as cancer, is of major interest in current molecular epidemiology. Nevertheless, the tremendous number of SNPs on the human genome requires computational methods for prioritizing SNPs according to their potentially deleterious effects to human health, and as such, for expediting genotyping and analysis. As of yet, little has been done to quantitatively assess the possible deleterious effects of SNPs for effective association studies. RESULTS: We propose a new integrative scoring system for prioritizing SNPs based on their possible deleterious effects within a probabilistic framework. We applied our system to 580 disease-susceptibility genes obtained from the OMIM (Online Mendelian Inheritance in Man) database, which is one of the most widely used databases of human genes and genetic disorders. The scoring results clearly show that the distribution of the functional significance (FS) scores for already known disease-related SNPs is significantly different from that of neutral SNPs. In addition, we summarize distinct features of potentially deleterious SNPs based on their FS score, such as functional genomic regions where they occur or bio-molecular functions that they mainly affect. We also demonstrate, through a comparative study, that our system improves upon other function-assessment systems for SNPs, by assigning significantly higher FS scores to already known disease-related SNPs than to neutral SNPs. Phil Hyoun Lee, Hagit Shatkay |
Bioinform. | 2 |
| 2009 | Functionally informative tag SNP selection using a pareto-optimal approach: playing the game of life
Phil Hyoun Lee, Hagit Shatkay |
BMC Bioinform. | 3 |
| 2009 | How to Get the Most out of Your Curation EffortabstractLarge-scale annotation efforts typically involve several experts who may disagree with each other. We propose an approach for modeling disagreements among experts that allows providing each annotation with a confidence value (i.e., the posterior probability that it is correct). Our approach allows computing certainty-level for individual annotations, given annotator-specific parameters estimated from data. We developed two probabilistic models for performing this analysis, compared these models using computer simulation, and tested each model's actual performance, based on a large data set generated by human annotators specifically for this study. We show that even in the worst-case scenario, when all annotators disagree, our approach allows us to significantly increase the probability of choosing the correct annotation. Along with this publication we make publicly available a corpus of 10,000 sentences annotated according to several cardinal dimensions that we have introduced in earlier work. The 10,000 sentences were all 3-fold annotated by a group of eight experts, while a 1,000-sentence subset was further 5-fold annotated by five new experts. While the presented data represent a specialized curation task, our modeling approach is general; most data annotation studies could benefit from our methodology. Andrey Rzhetsky, Hagit Shatkay, W. John Wilbur |
PLoS Comput. Biol. | 2 |
| 2008 | Ranking single nucleotide polymorphisms by potential deleterious effects
Phil Hyoun Lee, Hagit Shatkay |
AMIA | 2 |
| 2008 | Multi-dimensional classification of biomedical text: Toward automated, practical provision of high-utility text to diverse usersabstractMOTIVATION: Much current research in biomedical text mining is concerned with serving biologists by extracting certain information from scientific text. We note that there is no 'average biologist' client; different users have distinct needs. For instance, as noted in past evaluation efforts (BioCreative, TREC, KDD) database curators are often interested in sentences showing experimental evidence and methods. Conversely, lab scientists searching for known information about a protein may seek facts, typically stated with high confidence. Text-mining systems can target specific end-users and become more effective, if the system can first identify text regions rich in the type of scientific content that is of interest to the user, retrieve documents that have many such regions, and focus on fact extraction from these regions. Here, we study the ability to characterize and classify such text automatically. We have recently introduced a multi-dimensional categorization and annotation scheme, developed to be applicable to a wide variety of biomedical documents and scientific statements, while intended to support specific biomedical retrieval and extraction tasks. RESULTS: The annotation scheme was applied to a large corpus in a controlled effort by eight independent annotators, where three individual annotators independently tagged each sentence. We then trained and tested machine learning classifiers to automatically categorize sentence fragments based on the annotation. We discuss here the issues involved in this task, and present an overview of the results. The latter strongly suggest that automatic annotation along most of the dimensions is highly feasible, and that this new framework for scientific sentence categorization is applicable in practice. Hagit Shatkay, Fengxia Pan, Andrey Rzhetsky, W. John Wilbur |
Bioinform. | 1 |
| 2008 | Ranking single nucleotide polymorphisms by potential deleterious effects
Phil Hyoun Lee, Hagit Shatkay |
BMC Bioinform. | 2 |
| 2007 | Using Cluster Ensemble and Validation to Identify Subtypes of Pervasive Developmental Disorders
Jess J. Shen, Phil Hyoun Lee, Jeanette J. A. Holden, Hagit Shatkay |
AMIA | 4 |
| 2007 | Two Birds, One Stone: Selecting Functionally Informative Tag SNPs for Disease Association Studies
Phil Hyoun Lee, Hagit Shatkay |
WABI | 2 |
| 2007 | SherLoc: high-accuracy prediction of protein subcellular localization by integrating text and protein sequence dataabstractMOTIVATION: Knowing the localization of a protein within the cell helps elucidate its role in biological processes, its function and its potential as a drug target. Thus, subcellular localization prediction is an active research area. Numerous localization prediction systems are described in the literature; some focus on specific localizations or organisms, while others attempt to cover a wide range of localizations. RESULTS: We introduce SherLoc, a new comprehensive system for predicting the localization of eukaryotic proteins. It integrates several types of sequence and text-based features. While applying the widely used support vector machines (SVMs), SherLoc's main novelty lies in the way in which it selects its text sources and features, and integrates those with sequence-based features. We test SherLoc on previously used datasets, as well as on a new set devised specifically to test its predictive power, and show that SherLoc consistently improves on previous reported results. We also report the results of applying SherLoc to a large set of yet-unlocalized proteins. AVAILABILITY: SherLoc, along with Supplementary Information, is available at: http://www-bs.informatik.uni-tuebingen.de/Services/SherLoc/ Hagit Shatkay, Annette Höglund, Scott Brady, Torsten Blum, Pierre Dönnes, Oliver Kohlbacher |
Bioinform. | 1 |
| 2006 | Combining multi-species genomic data for microRNA identification using a Naïve Bayes classifierabstractMOTIVATION: Most computational methodologies for microRNA gene prediction utilize techniques based on sequence conservation and/or structural similarity. In this study we describe a new technique, which is applicable across several species, for predicting miRNA genes. This technique is based on machine learning, using the Naive Bayes classifier. It automatically generates a model from the training data, which consists of sequence and structure information of known miRNAs from a variety of species. RESULTS: Our study shows that the application of machine learning techniques, along with the integration of data from multiple species is a useful and general approach for miRNA gene prediction. Based on our experiments, we believe that this new technique is applicable to an extensive range of eukaryotes' genomes. Specific structure and sequence features are first used to identify miRNAs followed by a comparative analysis to decrease the number of false positives (FPs). The resulting algorithm exhibits higher specificity and similar sensitivity compared to currently used algorithms that rely on conserved genomic regions to decrease the rate of FPs. Malik Yousef, Michael Nebozhyn, Hagit Shatkay, Stathis Kanterakis, Louise C. Showe, Michael K. Showe |
Bioinform. | 3 |
| 2006 | Discovering semantic features in the literature: a foundation for building functional associationsabstractBACKGROUND: Experimental techniques such as DNA microarray, serial analysis of gene expression (SAGE) and mass spectrometry proteomics, among others, are generating large amounts of data related to genes and proteins at different levels. As in any other experimental approach, it is necessary to analyze these data in the context of previously known information about the biological entities under study. The literature is a particularly valuable source of information for experiment validation and interpretation. Therefore, the development of automated text mining tools to assist in such interpretation is one of the main challenges in current bioinformatics research. RESULTS: We present a method to create literature profiles for large sets of genes or proteins based on common semantic features extracted from a corpus of relevant documents. These profiles can be used to establish pair-wise similarities among genes, utilized in gene/protein classification or can be even combined with experimental measurements. Semantic features can be used by researchers to facilitate the understanding of the commonalities indicated by experimental results. Our approach is based on non-negative matrix factorization (NMF), a machine-learning algorithm for data analysis, capable of identifying local patterns that characterize a subset of the data. The literature is thus used to establish putative relationships among subsets of genes or proteins and to provide coherent justification for this clustering into subsets. We demonstrate the utility of the method by applying it to two independent and vastly different sets of genes. CONCLUSION: The presented method can create literature profiles from documents relevant to sets of genes. The representation of genes as additive linear combinations of semantic features allows for the exploration of functional associations as well as for clustering, suggesting a valuable methodology for the validation and interpretation of high-throughput experimental data. Monica Chagoyen, Pedro Carmona-Saez, Hagit Shatkay, José María Carazo, Alberto D. Pascual-Montano |
BMC Bioinform. | 3 |
| 2006 | New directions in biomedical text annotation: definitions, guidelines and corpus constructionabstractBACKGROUND: While biomedical text mining is emerging as an important research area, practical results have proven difficult to achieve. We believe that an important first step towards more accurate text-mining lies in the ability to identify and characterize text that satisfies various types of information needs. We report here the results of our inquiry into properties of scientific text that have sufficient generality to transcend the confines of a narrow subject area, while supporting practical mining of text for factual information. Our ultimate goal is to annotate a significant corpus of biomedical text and train machine learning methods to automatically categorize such text along certain dimensions that we have defined. RESULTS: We have identified five qualitative dimensions that we believe characterize a broad range of scientific sentences, and are therefore useful for supporting a general approach to text-mining: focus, polarity, certainty, evidence, and directionality. We define these dimensions and describe the guidelines we have developed for annotating text with regard to them. To examine the effectiveness of the guidelines, twelve annotators independently annotated the same set of 101 sentences that were randomly selected from current biomedical periodicals. Analysis of these annotations shows 70-80% inter-annotator agreement, suggesting that our guidelines indeed present a well-defined, executable and reproducible task. CONCLUSION: We present our guidelines defining a text annotation task, along with annotation results from multiple independently produced annotations, demonstrating the feasibility of the task. The annotation of a very large corpus of documents along these guidelines is currently ongoing. These annotations form the basis for the categorization of text along multiple dimensions, to support viable text mining for experimental results, methodology statements, and other forms of information. We are currently developing machine learning methods, to be trained and tested on the annotated corpus, that would allow for the automatic categorization of biomedical text along the general dimensions that we have presented. The guidelines in full detail, along with annotated examples, are publicly available. W. John Wilbur, Andrey Rzhetsky, Hagit Shatkay |
BMC Bioinform. | 3 |
| 2005 | Hairpins in bookstacks: Information retrieval from biomedical textabstractCurrent advances in high-throughput biology are accompanied by a tremendous increase in the number of related publications. Much biomedical information is reported in the vast amount of literature. The ability to rapidly and effectively survey the literature is necessary for both the design and the interpretation of large-scale experiments, and for curation of structured biomedical knowledge in public databases. Given the millions of published documents, the field of information retrieval, which is concerned with the automatic identification of relevant documents from large text collections, has much to offer. This paper introduces the basics of information retrieval, discusses its applications in biomedicine, and presents traditional and non-traditional ways in which it can be used. Hagit Shatkay |
Briefings Bioinform. | 1 |
| 2002 | Learning Geometrically-Constrained Hidden Markov Models for Robot Navigation: Bridging the Topological-Geometrical GapabstractHidden Markov models (HMMs) and partially observable Markov decision processes (POMDPs) provide useful tools for modeling dynamical systems. They are particularly useful for representing the topology of environments such as road networks and office buildings, which are typical for robot navigation and planning. The work presented here describes a formal framework for incorporating readily available odometric information and geometrical constraints into both the models and the algorithm that learns them. By taking advantage of such information, learning HMMs/POMDPs can be made to generate better solutions and require fewer iterations, while being robust in the face of data reduction. Experimental results, obtained from both simulated and real robot data, demonstrate the effectiveness of the approach. Hagit Shatkay, Leslie Pack Kaelbling |
J. Artif. Intell. Res. | 1 |
| 2000 | Genes, Themes, and Microarrays: Using Information Retrieval for Large-Scale Gene Analysis
Hagit Shatkay, Stephen Edwards, W. John Wilbur, Mark Boguski |
ISMB | 1 |
| 1999 | Learning Hidden Markov Models with Geometrical Constraints
Hagit Shatkay |
UAI | 1 |
| 1998 | Heading in the Right Direction
Hagit Shatkay, Leslie Pack Kaelbling |
ICML | 1 |
| 1997 | Learning Topological Maps with Weak Local Odometric Information
Hagit Shatkay, Leslie Pack Kaelbling |
IJCAI (2) | 1 |
| 1996 | Approximate Queries and Representations for Large Data SequencesabstractMany new database application domains such as experimental sciences and medicine are characterized by large sequences as their main form of data. Using approximate representation can significantly reduce the required storage and search space. A good choice of representation, can support a broad new class of approximate queries, needed in there domains. These queries are concerned with application dependent features of the data as opposed to the actual sampled points. We introduce a new notion of generalized approximate queries and a general divide and conquer approach that supports them. This approach uses families of real-valued functions as an approximate representation. We present an algorithm for realizing our technique, and the results of applying it to medical cardiology data. Hagit Shatkay, Stanley B. Zdonik |
ICDE | 1 |