EDBT 2026 Demo / reviewers in the wild / expert
Antonio Jimeno-Yepes
dblp:50/8088 · also Antonio José Jimeno-Yepes
· DBLP profile ↗
47ranked-venue papers
18as first author
10since 2021 · last 2025
0000-0002-6581-094XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 30 · 13 first-author · 5 since 2021Artificial intelligence and machine learning · 15 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-Label Generalized Zero Shot Chest X-Ray Classification by Combining Image-Text Information With Feature DisentanglementabstractIn fully supervised learning-based medical image classification, the robustness of a trained model is influenced by its exposure to the range of candidate disease classes. Generalized Zero Shot Learning (GZSL) aims to correctly predict seen and novel unseen classes. Current GZSL approaches have focused mostly on the single-label case. However, it is common for chest X-rays to be labelled with multiple disease classes. We propose a novel multi-modal multi-label GZSL approach that leverages feature disentanglement andmulti-modal information to synthesize features of unseen classes. Disease labels are processed through a pre-trained BioBert model to obtain text embeddings that are used to create a dictionary encoding similarity among different labels. We then use disentangled features and graph aggregation to learn a second dictionary of inter-label similarities. A subsequent clustering step helps to identify representative vectors for each class. The multi-modal multi-label dictionaries and the class representative vectors are used to guide the feature synthesis step, which is the most important component of our pipeline, for generating realistic multi-label disease samples of seen and unseen classes. Our method is benchmarked against multiple competing methods and we outperform all of them based on experiments conducted on the publicly available NIH and CheXpert chest X-ray datasets. Dwarikanath Mahapatra, Antonio Jimeno-Yepes, Behzad Bozorgtabar, Sudipta Roy 0002, ZongYuan Ge, Mauricio Reyes 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2025 | Corrections to "Multi-Label Generalized Zero Shot Chest X-Ray Classification By Combining Image-Text Information With Feature Disentanglement"abstractPresents corrections to the paper, (Corrections to "Multi-Label Generalized Zero Shot Chest X-Ray Classification By Combining Image-Text Information With Feature Disentanglement"). Dwarikanath Mahapatra, Antonio Jimeno-Yepes, Behzad Bozorgtabar, Sudipta Roy 0002, ZongYuan Ge, Mauricio Reyes 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Boosting Patient Representation Learning via Graph Contrastive Learning
Yuxi Liu 0003, Jiang Bian 0001, Antonio Jimeno-Yepes, Jun Shen 0001, Fuyi Li, Guodong Long, Flora D. Salim |
ECML/PKDD (9) | 4 |
| 2023 | Stacked Attention-based Networks for Accurate and Interpretable Health Risk PredictionabstractPredicting the health risks of patients based on electronic health records (EHRs) has recently attracted considerable research interest. Health risk refers to the probability of the occurrence of a specific health outcome for a specific patient. The predicted risks of a specific health outcome can be used to support decisions by healthcare professionals. Various predictive models have been developed. Compared with traditional machine learning models, deep learning-based models have achieved more promising performance. However, due to the lack of transparency, the acceptance of deep learning-based models are often limited. This paper proposes a Stacked Attention-based Network, SANet, for accurate and interpretable health risk prediction. Two novel attention-based modules, named Convolutional Attention Module and Sequential Attention Module respectively, are designed to capture patient-specific contextual information at both feature and sequence levels. Particularly, Sequential Attention Module can flexibly learn the impact of the time interval between sequential visits and significantly enhance the interpretability and robustness of learning outcomes from sequences. Experimental results on two real-world EHR datasets demonstrate the superior predictive accuracy of our method, as well as interpretability and robustness, compared to existing state-of-the-art methods. The findings extracted by this approach are also empirically confirmed by relevant literature and medical experts. Yuxi Liu 0003, Campbell Thompson, Richard Leibbrandt, Shaowen Qin, Antonio Jimeno-Yepes |
IJCNN | 6 |
| 2023 | Class Specific Feature Disentanglement and Text Embeddings for Multi-label Generalized Zero Shot CXR Classification
Dwarikanath Mahapatra, Antonio Jimeno-Yepes, Shiba Kuanar, Sudipta Roy 0002, Behzad Bozorgtabar, Mauricio Reyes 0001, ZongYuan Ge |
MICCAI (2) | 2 |
| 2022 | Integrated Convolutional and Recurrent Neural Networks for Health Risk Prediction using Patient Journey Data with Many Missing ValuesabstractPredicting the health risks of patients using Electronic Health Records (EHR) has attracted considerable attention in recent years, especially with the development of deep learning techniques. Health risk refers to the probability of the occurrence of a specific health outcome for a specific patient. The predicted risks can be used to support decision-making by healthcare professionals. EHRs are structured patient journey data. Each patient journey contains a chronological set of clinical events, and within each clinical event, there is a set of clinical/medical activities. Due to variations of patient conditions and treatment needs, EHR patient journey data has an inherently high degree of missingness that contains important information affecting relationships among variables, including time. Existing deep learning-based models generate imputed values for missing values when learning the relationships. However, imputed data in EHR patient journey data may distort the clinical meaning of the original EHR patient journey data, resulting in classification bias. This paper proposes a novel end-to-end approach to modeling EHR patient journey data with Integrated Convolutional and Recurrent Neural Networks. Our model can capture both long- and short-term temporal patterns within each patient journey and effectively handle the high degree of missingness in EHR data without any imputation data generation. Extensive experimental results using the proposed model on two real-world datasets demonstrate robust performance as well as superior prediction accuracy compared to existing state-of-the-art imputation-based prediction methods. Yuxi Liu 0003, Shaowen Qin, Antonio Jimeno-Yepes, Wei Shao 0006, Flora D. Salim |
BIBM | 3 |
| 2021 | Brief Description of COVID-SEE: The Scientific Evidence Explorer for COVID-19 Related Research
Karin Verspoor, Simon Suster, Yulia Otmakhova 0001, Shevon Mendis, Zenan Zhai, Biaoyan Fang, Jey Han Lau, Timothy Baldwin, Antonio Jimeno-Yepes, David Martínez 0001 |
ECIR (2) | 9 |
| 2021 | ICDAR 2021 Competition on Scientific Literature Parsing
Antonio Jimeno-Yepes, Peter Zhong, Douglas Burdick |
ICDAR (4) | 1 |
| 2021 | Grey-box Adversarial Attack And Defence For Sentiment ClassificationabstractYing Xu, Xu Zhong, Antonio Jimeno Yepes, Jey Han Lau. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Xu Zhong, Antonio Jimeno-Yepes, Jey Han Lau |
NAACL-HLT | 3 |
| 2021 | A representation model for biological entities by fusing structured axioms with unstructured textsabstractMOTIVATION: Structured semantic resources, for example, biological knowledge bases and ontologies, formally define biological concepts, entities and their semantic relationships, manifested as structured axioms and unstructured texts (e.g. textual definitions). The resources contain accurate expressions of biological reality and have been used by machine-learning models to assist intelligent applications like knowledge discovery. The current methods use both the axioms and definitions as plain texts in representation learning (RL). However, since the axioms are machine-readable while the natural language is human-understandable, difference in meaning of token and structure impedes the representations to encode desirable biological knowledge. RESULTS: We propose ERBK, a RL model of bio-entities. Instead of using the axioms and definitions as a textual corpus, our method uses knowledge graph embedding method and deep convolutional neural models to encode the axioms and definitions respectively. The representations could not only encode more underlying biological knowledge but also be further applied to zero-shot circumstance where existing approaches fall short. Experimental evaluations show that ERBK outperforms the existing methods for predicting protein-protein interactions and gene-disease associations. Moreover, it shows that ERBK still maintains promising performance under the zero-shot circumstance. We believe the representations and the method have certain generality and could extend to other types of bio-relation. AVAILABILITY AND IMPLEMENTATION: The source code is available at the gitlab repository https://gitlab.com/BioAI/erbk. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Peiliang Lou, Yuxin Dong 0003, Antonio Jimeno-Yepes, Chen Li 0011 |
Bioinform. | 3 |
| 2020 | Prediction of secondary structure population and intrinsic disorder of proteins using multitask deep learning
André Leier, Tatiana T. Marquez-Lago, Jue Xie, Antonio Jimeno-Yepes, James C. Whisstock, Campbell Wilson, Jiangning Song |
AMIA | 5 |
| 2020 | Image-Based Table Recognition: Data, Model, and Evaluation
Xu Zhong, Elaheh ShafieiBavani, Antonio Jimeno-Yepes |
ECCV (21) | 3 |
| 2020 | Forget Me Not: Reducing Catastrophic Forgetting for Domain Adaptation in Reading ComprehensionabstractThe creation of large-scale open domain reading comprehension data sets in recent years has enabled the development of end-to-end neural comprehension models with promising results. To use these models for domains with limited training data, one of the most effective approach is to first pre-train them on large out-of-domain source data and then fine-tune them with the limited target data. The caveat of this is that after fine-tuning the comprehension models tend to perform poorly in the source domain, a phenomenon known as catastrophic forgetting. In this paper, we explore methods that reduce catastrophic forgetting during fine-tuning without assuming access to data from the source domain. We introduce new auxiliary penalty terms and observe the best performance when a combination of auxiliary penalty terms is used to regularise the fine-tuning process for adapting comprehension models. To test our methods, we develop and release 6 narrow domain data sets that can potentially be used as reading comprehension benchmarks. Xu Zhong, Antonio Jimeno-Yepes, Jey Han Lau |
IJCNN | 3 |
| 2020 | MEDLINE as a Parallel Corpus: a Survey to Gain Insight on French-, Spanish- and Portuguese-speaking Authors' Abstract Writing PracticeabstractBackground: Parallel corpora are used to train and evaluate machine translation systems. To alleviate the cost of producing parallel resources for evaluation campaigns, existing corpora are leveraged. However, little information may be available about the methods used for producing the corpus, including translation direction. Objective: To gain insight on MEDLINE parallel corpus used in the biomedical task at the Workshop on Machine Translation in 2019 (WMT 2019). Material and Methods: Contact information for the authors of MEDLINE articles included in the English/Spanish (EN/ES), English/French (EN/FR), and English/Portuguese (EN/PT) WMT 2019 test sets was obtained from PubMed and publisher websites. The authors were asked about their abstract writing practices in a survey. Results: The response rate was above 20%. Authors reported that they are mainly native speakers of languages other than English. Although manual translation, sometimes via professional translation services, was commonly used for abstract translation, authors of articles in the EN/ES and EN/PT sets also relied on post-edited machine translation. Discussion: This study provides a characterization of MEDLINE authors’ language skills and abstract writing practices. Conclusion: The information collected in this study will be used to inform test set design for the next WMT biomedical task. Aurélie Névéol, Antonio Jimeno-Yepes, Mariana L. Neves |
LREC | 2 |
| 2020 | BioNorm: deep learning-based event normalization for the curation of reaction databasesabstractMOTIVATION: A biochemical reaction, bio-event, depicts the relationships between participating entities. Current text mining research has been focusing on identifying bio-events from scientific literature. However, rare efforts have been dedicated to normalize bio-events extracted from scientific literature with the entries in the curated reaction databases, which could disambiguate the events and further support interconnecting events into biologically meaningful and complete networks. RESULTS: In this paper, we propose BioNorm, a novel method of normalizing bio-events extracted from scientific literature to entries in the bio-molecular reaction database, e.g. IntAct. BioNorm considers event normalization as a paraphrase identification problem. It represents an entry as a natural language statement by combining multiple types of information contained in it. Then, it predicts the semantic similarity between the natural language statement and the statements mentioning events in scientific literature using a long short-term memory recurrent neural network (LSTM). An event will be normalized to the entry if the two statements are paraphrase. To the best of our knowledge, this is the first attempt of event normalization in the biomedical text mining. The experiments have been conducted using the molecular interaction data from IntAct. The results demonstrate that the method could achieve F-score of 0.87 in normalizing event-containing statements. AVAILABILITY AND IMPLEMENTATION: The source code is available at the gitlab repository https://gitlab.com/BioAI/leen and BioASQvec Plus is available on figshare https://figshare.com/s/45896c31d10c3f6d857a. Peiliang Lou, Antonio Jimeno-Yepes, Zai Zhang 0002, Xiangrong Zhang, Chen Li 0011 |
Bioinform. | 2 |
| 2020 | Adverse drug event detection using reason assignments in FDA drug labels
Corey Sutphin, Kahyun Lee, Antonio Jimeno-Yepes, Özlem Uzuner, Bridget T. McInnes |
J. Biomed. Informatics | 3 |
| 2019 | Neural Relation Extraction from Biomedical Literature
Elaheh ShafieiBavani, Antonio Jimeno-Yepes |
AMIA | 2 |
| 2019 | PubLayNet: Largest Dataset Ever for Document Layout AnalysisabstractRecognizing the layout of unstructured digital documents is an important step when parsing the documents into structured machine-readable format for downstream applications. Deep neural networks that are developed for computer vision have been proven to be an effective method to analyze layout of document images. However, document layout datasets that are currently publicly available are several magnitudes smaller than established computing vision datasets. Models have to be trained by transfer learning from a base model that is pre-trained on a traditional computer vision dataset. In this paper, we develop the PubLayNet dataset for document layout analysis by automatically matching the XML representations and the content of over 1 million PDF articles that are publicly available on PubMed Central. The size of the dataset is comparable to established computer vision datasets, containing over 360 thousand document images, where typical document layout elements are annotated. The experiments demonstrate that deep neural networks trained on PubLayNet accurately recognize the layout of scientific articles. The pre-trained models are also a more effective base mode for transfer learning on a different document domain. We release the dataset (https://github.com/ibm-aur-nlp/PubLayNet) to support development and evaluation of more advanced models for document layout analysis. Xu Zhong, Jianbin Tang, Antonio Jimeno-Yepes |
ICDAR | 3 |
| 2018 | A hybrid approach for automated mutation annotation of the extended human mutation landscape in scientific literature
Antonio Jimeno-Yepes, Andrew MacKinlay, Natalie Gunn, Christine Schieber, Noel Faux, Matthew Downton, Benjamin Goudey |
AMIA | 1 |
| 2018 | Parallel Corpora for the Biomedical Domain
Aurélie Névéol, Antonio Jimeno-Yepes, Mariana L. Neves, Karin Verspoor |
LREC | 2 |
| 2018 | Semantic Labeling Using a Low-Power Neuromorphic PlatformabstractDeep learning is a powerful technique for the analysis of remote sensing imagery. For applications that require real-time processing on mobile platforms, a low power consumption processing unit is advantageous. The human brain is remarkably powerful at image recognition tasks while operating at very low power consumption levels. Neuromorphic computing designs aim to achieve energy efficiency through the use of spiking neurons and low-precision synapses to perform data processing. We demonstrate here the classification of red, green, blue and depth and hyperspectral data sets using a neuromorphic processing unit (IBM TrueNorth Neurosynaptic System). The convolutional neural-network architecture of the classifier network has been adapted to fit the neuromorphic architecture. The results on overhead imagery and hyperspectral imagery data show that neuromorphic platforms can achieve the state-of-theart performance in semantic labeling with significantly (≈1000×) lower power consumption than traditional GPU-based solutions. Jianbin Tang, Benjamin S. Mashford, Antonio Jimeno-Yepes |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2017 | Detection of adverse drug reactions using medical named entities on Twitter
Andrew MacKinlay, Hafsah Aamer, Antonio Jimeno-Yepes |
AMIA | 3 |
| 2017 | Knowledge-Based Biomedical Word Sense Disambiguation with Neural Concept EmbeddingsabstractBiomedical word sense disambiguation (WSD) is an important intermediate task in many natural language processing applications such as named entity recognition, syntactic parsing, and relation extraction. In this paper, we employ knowledge-based approaches that also exploit recent advances in neural word/concept embeddings to improve over the state-of-the-art in biomedical WSD using the public MSH WSD dataset [1] as the test set. Our methods involve weak supervision - we do not use any hand-labeled examples for WSD to build our prediction models; however, we employ an existing concept mapping program, MetaMap, to obtain our concept vectors. Over the MSH WSD dataset, our linear time (in terms of numbers of senses and words in the test instance) method achieves an accuracy of 92.24% which is a 3% improvement over the best known results [2] obtained via unsupervised means. A more expensive approach that we developed relies on a nearest neighbor framework and achieves accuracy of 94.34%, essentially cutting the error rate in half. Employing dense vector representations learned from unlabeled free text has been shown to benefit many language processing tasks recently and our efforts show that biomedical WSD is no exception to this trend. For a complex and rapidly evolving domain such as biomedicine, building labeled datasets for larger sets of ambiguous terms may be impractical. Here, we show that weak supervision that leverages recent advances in representation learning can rival supervised approaches in biomedical WSD. However, external knowledge bases (here sense inventories) play a key role in the improvements achieved. Akm Sabbir, Antonio Jimeno-Yepes, Ramakanth Kavuluru |
BIBE | 2 |
| 2017 | Improving Classification Accuracy of Feedforward Neural Networks for Spiking Neuromorphic ChipsabstractDeep Neural Networks (DNN) achieve human level performance in many image analytics tasks but DNNs are mostly deployed to GPU platforms that consume a considerable amount of power. New hardware platforms using lower precision arithmetic achieve drastic reductions in power consumption. More recently, brain-inspired spiking neuromorphic chips have achieved even lower power consumption, on the order of milliwatts, while still offering real-time processing. However, for deploying DNNs to energy efficient neuromorphic chips the incompatibility between continuous neurons and synaptic weights of traditional DNNs, discrete spiking neurons and synapses of neuromorphic chips need to be overcome. Previous work has achieved this by training a network to learn continuous probabilities, before it is deployed to a neuromorphic architecture, such as IBM TrueNorth Neurosynaptic System, by random sampling these probabilities. The main contribution of this paper is a new learning algorithm that learns a TrueNorth configuration ready for deployment. We achieve this by training directly a binary hardware crossbar that accommodates the TrueNorth axon configuration constrains and we propose a different neuron model. Results of our approach trained on electroencephalogram (EEG) data show a significant improvement with previous work (76% vs 86% accuracy) while maintaining state of the art performance on the MNIST handwritten data set. Antonio Jimeno-Yepes, Jianbin Tang, Benjamin S. Mashford |
IJCAI | 1 |
| 2017 | Named Entity Recognition with Stack Residual LSTM and Trainable Bias DecodingabstractRecurrent Neural Network models are the state-of-the-art for Named Entity Recognition (NER). We present two innovations to improve the performance of these models. The first innovation is the introduction of residual connections between the Stacked Recurrent Neural Network model to address the degradation problem of deep neural networks. The second innovation is a bias decoding mechanism that allows the trained system to adapt to non-differentiable and externally computed objectives, such as the entity-based F-measure. Our work improves the state-of-the-art results for both Spanish and English languages on the standard train/development/test split of the CoNLL 2003 Shared Task NER dataset. Quan Tran, Andrew MacKinlay, Antonio Jimeno-Yepes |
IJCNLP(1) | 3 |
| 2017 | Word embeddings and recurrent neural networks based on Long-Short Term Memory nodes in supervised biomedical word sense disambiguation
Antonio Jimeno-Yepes |
J. Biomed. Informatics | 1 |
| 2016 | Weighted Population Code for low power neuromorphic image classificationabstractRecent digital spiking neuromorphic chips can perform complex computations in real-time with very low power consumption. The input data to such systems needs to first be converted into spikes using a spike encoding scheme. Current examples of such schemes include rate codes and population codes. The selected coding scheme might heavily impact the system's energy consumption, communication bandwidth, processing frame-rate, and computation accuracy. Hence it is important to make an educated decision when selecting the most appropriate spike coding scheme for a given task. To this end, we present a novel spike coding scheme named Weighted Population Code (WPC). WPC is compared to existing coding schemes to transduce images for classification using the TrueNorth chip. Extensive on-chip experimentation with the MNIST and the Flickr-LOGOS32 datasets sheds light on the trade-offs between accuracy, bandwidth, frame rate, network size and energy consumption for image classification, showing the advantages of WPC when high dynamic range and accuracy are needed. Antonio Jimeno-Yepes, Jianbin Tang, Shreya Saxena, Tobias Brosch, Arnon Amir |
IJCNN | 1 |
| 2016 | The Scielo Corpus: a Parallel Corpus of Scientific Publications for Biomedicine
Mariana L. Neves, Antonio Jimeno-Yepes, Aurélie Névéol |
LREC | 2 |
| 2015 | Feature engineering for MEDLINE citation categorization with MeSHabstractBACKGROUND: Research in biomedical text categorization has mostly used the bag-of-words representation. Other more sophisticated representations of text based on syntactic, semantic and argumentative properties have been less studied. In this paper, we evaluate the impact of different text representations of biomedical texts as features for reproducing the MeSH annotations of some of the most frequent MeSH headings. In addition to unigrams and bigrams, these features include noun phrases, citation meta-data, citation structure, and semantic annotation of the citations. RESULTS: Traditional features like unigrams and bigrams exhibit strong performance compared to other feature sets. Little or no improvement is obtained when using meta-data or citation structure. Noun phrases are too sparse and thus have lower performance compared to more traditional features. Conceptual annotation of the texts by MetaMap shows similar performance compared to unigrams, but adding concepts from the UMLS taxonomy does not improve the performance of using only mapped concepts. The combination of all the features performs largely better than any individual feature set considered. In addition, this combination improves the performance of a state-of-the-art MeSH indexer. Concerning the machine learning algorithms, we find that those that are more resilient to class imbalance largely obtain better performance. CONCLUSIONS: We conclude that even though traditional features such as unigrams and bigrams have strong performance compared to other features, it is possible to combine them to effectively improve the performance of the bag-of-words representation. We have also found that the combination of the learning algorithm and feature sets has an influence in the overall performance of the system. Moreover, using learning algorithms resilient to class imbalance largely improves performance. However, when using a large set of features, consideration needs to be taken with algorithms due to the risk of over-fitting. Specific combinations of learning algorithms and features for individual MeSH headings could further increase the performance of an indexing system. Antonio Jimeno-Yepes, Laura Plaza, Jorge Carrillo de Albornoz, James G. Mork, Alan R. Aronson |
BMC Bioinform. | 1 |
| 2015 | Knowledge based word-concept model estimation and refinement for biomedical text mining
Antonio Jimeno-Yepes, Rafael Berlanga Llavori |
J. Biomed. Informatics | 1 |
| 2013 | Comparison and combination of several MeSH indexing approaches
Antonio Jimeno-Yepes, James G. Mork, Dina Demner-Fushman, Alan R. Aronson |
AMIA | 1 |
| 2013 | MeSH indexing based on automatically generated summariesabstractBACKGROUND: MEDLINE citations are manually indexed at the U.S. National Library of Medicine (NLM) using as reference the Medical Subject Headings (MeSH) controlled vocabulary. For this task, the human indexers read the full text of the article. Due to the growth of MEDLINE, the NLM Indexing Initiative explores indexing methodologies that can support the task of the indexers. Medical Text Indexer (MTI) is a tool developed by the NLM Indexing Initiative to provide MeSH indexing recommendations to indexers. Currently, the input to MTI is MEDLINE citations, title and abstract only. Previous work has shown that using full text as input to MTI increases recall, but decreases precision sharply. We propose using summaries generated automatically from the full text for the input to MTI to use in the task of suggesting MeSH headings to indexers. Summaries distill the most salient information from the full text, which might increase the coverage of automatic indexing approaches based on MEDLINE. We hypothesize that if the results were good enough, manual indexers could possibly use automatic summaries instead of the full texts, along with the recommendations of MTI, to speed up the process while maintaining high quality of indexing results. RESULTS: We have generated summaries of different lengths using two different summarizers, and evaluated the MTI indexing on the summaries using different algorithms: MTI, individual MTI components, and machine learning. The results are compared to those of full text articles and MEDLINE citations. Our results show that automatically generated summaries achieve similar recall but higher precision compared to full text articles. Compared to MEDLINE citations, summaries achieve higher recall but lower precision. CONCLUSIONS: Our results show that automatic summaries produce better indexing than full text articles. Summaries produce similar recall to full text but much better precision, which seems to indicate that automatic summaries can efficiently capture the most important contents within the original articles. The combination of MEDLINE citations and automatically generated summaries could improve the recommendations suggested by MTI. On the other hand, indexing performance might be dependent on the MeSH heading being indexed. Summarization techniques could thus be considered as a feature selection algorithm that might have to be tuned individually for each MeSH heading. Antonio Jimeno-Yepes, Laura Plaza, James G. Mork, Alan R. Aronson, Alberto Díaz 0001 |
BMC Bioinform. | 1 |
| 2013 | Combining MEDLINE and publisher data to create parallel corpora for the automatic translation of biomedical textabstractBACKGROUND: Most of the institutional and research information in the biomedical domain is available in the form of English text. Even in countries where English is an official language, such as the United States, language can be a barrier for accessing biomedical information for non-native speakers. Recent progress in machine translation suggests that this technique could help make English texts accessible to speakers of other languages. However, the lack of adequate specialized corpora needed to train statistical models currently limits the quality of automatic translations in the biomedical domain. RESULTS: We show how a large-sized parallel corpus can automatically be obtained for the biomedical domain, using the MEDLINE database. The corpus generated in this work comprises article titles obtained from MEDLINE and abstract text automatically retrieved from journal websites, which substantially extends the corpora used in previous work. After assessing the quality of the corpus for two language pairs (English/French and English/Spanish) we use the Moses package to train a statistical machine translation model that outperforms previous models for automatic translation of biomedical text. CONCLUSIONS: We have built translation data sets in the biomedical domain that can easily be extended to other languages available in MEDLINE. These sets can successfully be applied to train statistical machine translation models. While further progress should be made by incorporating out-of-domain corpora and domain-specific lexicons, we believe that this work improves the automatic translation of biomedical texts. Antonio Jimeno-Yepes, Élise Prieur, Aurélie Névéol |
BMC Bioinform. | 1 |
| 2013 | GeneRIF indexing: sentence selection based on machine learningabstractBACKGROUND: A Gene Reference Into Function (GeneRIF) describes novel functionality of genes. GeneRIFs are available from the National Center for Biotechnology Information (NCBI) Gene database. GeneRIF indexing is performed manually, and the intention of our work is to provide methods to support creating the GeneRIF entries. The creation of GeneRIF entries involves the identification of the genes mentioned in MEDLINE®; citations and the sentences describing a novel function. RESULTS: We have compared several learning algorithms and several features extracted or derived from MEDLINE sentences to determine if a sentence should be selected for GeneRIF indexing. Features are derived from the sentences or using mechanisms to augment the information provided by them: assigning a discourse label using a previously trained model, for example. We show that machine learning approaches with specific feature combinations achieve results close to one of the annotators. We have evaluated different feature sets and learning algorithms. In particular, Naïve Bayes achieves better performance with a selection of features similar to one used in related work, which considers the location of the sentence, the discourse of the sentence and the functional terminology in it. CONCLUSIONS: The current performance is at a level similar to human annotation and it shows that machine learning can be used to automate the task of sentence selection for GeneRIF annotation. The current experiments are limited to the human species. We would like to see how the methodology can be extended to other species, specifically the normalization of gene mentions in other species. Antonio Jimeno-Yepes, J. Caitlin Sticco, James G. Mork, Alan R. Aronson |
BMC Bioinform. | 1 |
| 2011 | Exploiting MeSH indexing in MEDLINE to generate a data set for word sense disambiguationabstractBACKGROUND: Evaluation of Word Sense Disambiguation (WSD) methods in the biomedical domain is difficult because the available resources are either too small or too focused on specific types of entities (e.g. diseases or genes). We present a method that can be used to automatically develop a WSD test collection using the Unified Medical Language System (UMLS) Metathesaurus and the manual MeSH indexing of MEDLINE. We demonstrate the use of this method by developing such a data set, called MSH WSD. METHODS: In our method, the Metathesaurus is first screened to identify ambiguous terms whose possible senses consist of two or more MeSH headings. We then use each ambiguous term and its corresponding MeSH heading to extract MEDLINE citations where the term and only one of the MeSH headings co-occur. The term found in the MEDLINE citation is automatically assigned the UMLS CUI linked to the MeSH heading. Each instance has been assigned a UMLS Concept Unique Identifier (CUI). We compare the characteristics of the MSH WSD data set to the previously existing NLM WSD data set. RESULTS: The resulting MSH WSD data set consists of 106 ambiguous abbreviations, 88 ambiguous terms and 9 which are a combination of both, for a total of 203 ambiguous entities. For each ambiguous term/abbreviation, the data set contains a maximum of 100 instances per sense obtained from MEDLINE.We evaluated the reliability of the MSH WSD data set using existing knowledge-based methods and compared their performance to that of the results previously obtained by these algorithms on the pre-existing data set, NLM WSD. We show that the knowledge-based methods achieve different results but keep their relative performance except for the Journal Descriptor Indexing (JDI) method, whose performance is below the other methods. CONCLUSIONS: The MSH WSD data set allows the evaluation of WSD algorithms in the biomedical domain. Compared to previously existing data sets, MSH WSD contains a larger number of biomedical terms/abbreviations and covers the largest set of UMLS Semantic Types. Furthermore, the MSH WSD data set has been generated automatically reusing already existing annotations and, therefore, can be regenerated from subsequent UMLS versions. Antonio Jimeno-Yepes, Bridget T. McInnes, Alan R. Aronson |
BMC Bioinform. | 1 |
| 2011 | Collocation analysis for UMLS knowledge-based word sense disambiguationabstractBACKGROUND: The effectiveness of knowledge-based word sense disambiguation (WSD) approaches depends in part on the information available in the reference knowledge resource. Off the shelf, these resources are not optimized for WSD and might lack terms to model the context properly. In addition, they might include noisy terms which contribute to false positives in the disambiguation results. METHODS: We analyzed some collocation types which could improve the performance of knowledge-based disambiguation methods. Collocations are obtained by extracting candidate collocations from MEDLINE and then assigning them to one of the senses of an ambiguous word. We performed this assignment either using semantic group profiles or a knowledge-based disambiguation method. In addition to collocations, we used second-order features from a previously implemented approach.Specifically, we measured the effect of these collocations in two knowledge-based WSD methods. The first method, AEC, uses the knowledge from the UMLS to collect examples from MEDLINE which are used to train a Naïve Bayes approach. The second method, MRD, builds a profile for each candidate sense based on the UMLS and compares the profile to the context of the ambiguous word.We have used two WSD test sets which contain disambiguation cases which are mapped to UMLS concepts. The first one, the NLM WSD set, was developed manually by several domain experts and contains words with high frequency occurrence in MEDLINE. The second one, the MSH WSD set, was developed automatically using the MeSH indexing in MEDLINE. It contains a larger set of words and covers a larger number of UMLS semantic types. RESULTS: The results indicate an improvement after the use of collocations, although the approaches have different performance depending on the data set. In the NLM WSD set, the improvement is larger for the MRD disambiguation method using second-order features. Assignment of collocations to a candidate sense based on UMLS semantic group profiles is more effective in the AEC method.In the MSH WSD set, the increment in performance is modest for all the methods. Collocations combined with the MRD disambiguation method have the best performance. The MRD disambiguation method and second-order features provide an insignificant change in performance. The AEC disambiguation method gives a modest improvement in performance. Assignment of collocations to a candidate sense based on knowledge-based methods has better performance. CONCLUSIONS: Collocations improve the performance of knowledge-based disambiguation methods, although results vary depending on the test set and method used. Generally, the AEC method is sensitive to query drift. Using AEC, just a few selected terms provide a large improvement in disambiguation performance. The MRD method handles noisy terms better but requires a larger set of terms to improve performance. Antonio Jimeno-Yepes, Bridget T. McInnes, Alan R. Aronson |
BMC Bioinform. | 1 |
| 2011 | Studying the correlation between different word sense disambiguation methods and summarization effectiveness in biomedical textsabstractBACKGROUND: Word sense disambiguation (WSD) attempts to solve lexical ambiguities by identifying the correct meaning of a word based on its context. WSD has been demonstrated to be an important step in knowledge-based approaches to automatic summarization. However, the correlation between the accuracy of the WSD methods and the summarization performance has never been studied. RESULTS: We present three existing knowledge-based WSD approaches and a graph-based summarizer. Both the WSD approaches and the summarizer employ the Unified Medical Language System (UMLS) Metathesaurus as the knowledge source. We first evaluate WSD directly, by comparing the prediction of the WSD methods to two reference sets: the NLM WSD dataset and the MSH WSD collection. We next apply the different WSD methods as part of the summarizer, to map documents onto concepts in the UMLS Metathesaurus, and evaluate the summaries that are generated. The results obtained by the different methods in both evaluations are studied and compared. CONCLUSIONS: It has been found that the use of WSD techniques has a positive impact on the results of our graph-based summarizer, and that, when both the WSD and summarization tasks are assessed over large and homogeneous evaluation collections, there exists a correlation between the overall results of the WSD and summarization tasks. Furthermore, the best WSD algorithm in the first task tends to be also the best one in the second. However, we also found that the improvement achieved by the summarizer is not directly correlated with the WSD performance. The most likely reason is that the errors in disambiguation are not equally important but depend on the relative salience of the different concepts in the document to be summarized. Laura Plaza, Antonio Jimeno-Yepes, Alberto Díaz 0001, Alan R. Aronson |
BMC Bioinform. | 2 |
| 2010 | Query Expansion for UMLS Metathesaurus Disambiguation Based on Automatic Corpus ExtractionabstractWord sense disambiguation (WSD) is an intermediate task within information retrieval and information extraction, which attempts selecting the proper sense of ambiguous terms. In the biomedical domain, general WSD has not received much attention compared to the disambiguation of specific categories of entities like proteins and genes or diseases. Statistical learning approaches have achieved better performance compared to other methods. On the other hand, manually annotated data is limited, and covering all the ambiguous cases of a large resource like the UMLS is infeasible. Knowledge-based approaches using the UMLS and MEDLINE citations have achieved good performance but below that of statistical learning approaches. Our best knowledge-based result has been obtained by training a Naïve Bayes algorithm on an automatically extracted MEDLINE corpus. In this work, we extend on previous methods to enhance the quality of an automatically extracted corpus using related terms obtained from MEDLINE without manually annotated training data. We have focused on the extraction of collocations which might be used in combination with one of the senses of the ambiguous terms. We find that left side collocations have the largest improvement in accuracy with an improvement of 4%. In addition, the combination of different types of collocations and post-filtering of retrieved citations achieves an improvement of almost 9% in accuracy. Antonio Jimeno-Yepes, Alan R. Aronson |
ICMLA | 1 |
| 2010 | The CALBC Silver Standard Corpus for Biomedical Named Entities - A Study in Harmonizing the Contributions from Four Independent Named Entity Taggers
Dietrich Rebholz-Schuhmann, Antonio Jimeno-Yepes, Erik M. van Mulligen, Ning Kang 0002, Jan A. Kors, David Milward, Peter T. Corbett, Ekaterina Buyko, Katrin Tomanek, Elena Beisswanger, Udo Hahn |
LREC | 2 |
| 2010 | Knowledge-based biomedical word sense disambiguation: comparison of approachesabstractBACKGROUND: Word sense disambiguation (WSD) algorithms attempt to select the proper sense of ambiguous terms in text. Resources like the UMLS provide a reference thesaurus to be used to annotate the biomedical literature. Statistical learning approaches have produced good results, but the size of the UMLS makes the production of training data infeasible to cover all the domain. METHODS: We present research on existing WSD approaches based on knowledge bases, which complement the studies performed on statistical learning. We compare four approaches which rely on the UMLS Metathesaurus as the source of knowledge. The first approach compares the overlap of the context of the ambiguous word to the candidate senses based on a representation built out of the definitions, synonyms and related terms. The second approach collects training data for each of the candidate senses to perform WSD based on queries built using monosemous synonyms and related terms. These queries are used to retrieve MEDLINE citations. Then, a machine learning approach is trained on this corpus. The third approach is a graph-based method which exploits the structure of the Metathesaurus network of relations to perform unsupervised WSD. This approach ranks nodes in the graph according to their relative structural importance. The last approach uses the semantic types assigned to the concepts in the Metathesaurus to perform WSD. The context of the ambiguous word and semantic types of the candidate concepts are mapped to Journal Descriptors. These mappings are compared to decide among the candidate concepts. Results are provided estimating accuracy of the different methods on the WSD test collection available from the NLM. CONCLUSIONS: We have found that the last approach achieves better results compared to the other methods. The graph-based approach, using the structure of the Metathesaurus network to estimate the relevance of the Metathesaurus concepts, does not perform well compared to the first two methods. In addition, the combination of methods improves the performance over the individual approaches. On the other hand, the performance is still below statistical learning trained on manually produced data and below the maximum frequency sense baseline. Finally, we propose several directions to improve the existing methods and to improve the Metathesaurus to be more effective in WSD. Antonio Jimeno-Yepes, Alan R. Aronson |
BMC Bioinform. | 1 |
| 2010 | Ontology refinement for improved information retrieval
Antonio Jimeno-Yepes, Rafael Berlanga Llavori, Dietrich Rebholz-Schuhmann |
Inf. Process. Manag. | 1 |
| 2010 | Measuring prediction capacity of individual verbs for the identification of protein interactions
Dietrich Rebholz-Schuhmann, Antonio Jimeno-Yepes, Miguel Arregui, Harald Kirsch |
J. Biomed. Informatics | 2 |
| 2009 | Reuse of terminological resources for efficient ontological engineering in Life SciencesabstractThis paper is intended to explore how to use terminological resources for ontology engineering. Nowadays there are several biomedical ontologies describing overlapping domains, but there is not a clear correspondence between the concepts that are supposed to be equivalent or just similar. These resources are quite precious but their integration and further development are expensive. Terminologies may support the ontological development in several stages of the lifecycle of the ontology; e.g. ontology integration. In this paper we investigate the use of terminological resources during the ontology lifecycle. We claim that the proper creation and use of a shared thesaurus is a cornerstone for the successful application of the Semantic Web technology within life sciences. Moreover, we have applied our approach to a real scenario, the Health-e-Child (HeC) project, and we have evaluated the impact of filtering and re-organizing several resources. As a result, we have created a reference thesaurus for this project, named HeCTh. Antonio Jimeno-Yepes, Ernesto Jiménez-Ruiz, Rafael Berlanga Llavori, Dietrich Rebholz-Schuhmann |
BMC Bioinform. | 1 |
| 2009 | Annotation of protein residues based on a literature analysis: cross-validation against UniProtKbabstractBACKGROUND: A protein annotation database, such as the Universal Protein Resource knowledge base (UniProtKb), is a valuable resource for the validation and interpretation of predicted 3D structure patterns in proteins. Existing studies have focussed on point mutation extraction methods from biomedical literature which can be used to support the time consuming work of manual database curation. However, these methods were limited to point mutation extraction and do not extract features for the annotation of proteins at the residue level. RESULTS: This work introduces a system that identifies protein residues in MEDLINE abstracts and annotates them with features extracted from the context written in the surrounding text. MEDLINE abstract texts have been processed to identify protein mentions in combination with taxonomic species and protein residues (F1-measure 0.52). The identified protein-species-residue triplets have been validated and benchmarked against reference data resources (UniProtKb, average F1-measure of 0.54). Then, contextual features were extracted through shallow and deep parsing and the features have been classified into predefined categories (F1-measure ranges from 0.15 to 0.67). Furthermore, the feature sets have been aligned with annotation types in UniProtKb to assess the relevance of the annotations for ongoing curation projects. Altogether, the annotations have been assessed automatically and manually against reference data resources. CONCLUSION: This work proposes a solution for the automatic extraction of functional annotation for protein residues from biomedical articles. The presented approach is an extension to other existing systems in that a wider range of residue entities are considered and that features of residues are extracted as annotations. Kevin Nagel, Antonio Jimeno-Yepes, Dietrich Rebholz-Schuhmann |
BMC Bioinform. | 2 |
| 2008 | Text processing through Web services: calling WhatizitabstractMOTIVATION: Text-mining (TM) solutions are developing into efficient services to researchers in the biomedical research community. Such solutions have to scale with the growing number and size of resources (e.g. available controlled vocabularies), with the amount of literature to be processed (e.g. about 17 million documents in PubMed) and with the demands of the user community (e.g. different methods for fact extraction). These demands motivated the development of a server-based solution for literature analysis. Whatizit is a suite of modules that analyse text for contained information, e.g. any scientific publication or Medline abstracts. Special modules identify terms and then link them to the corresponding entries in bioinformatics databases such as UniProtKb/Swiss-Prot data entries and gene ontology concepts. Other modules identify a set of selected annotation types like the set produced by the EBIMed analysis pipeline for proteins. In the case of Medline abstracts, Whatizit offers access to EBI's in-house installation via PMID or term query. For large quantities of the user's own text, the server can be operated in a streaming mode (http://www.ebi.ac.uk/webservices/whatizit). Dietrich Rebholz-Schuhmann, Miguel Arregui, Sylvain Gaudan, Harald Kirsch, Antonio Jimeno-Yepes |
Bioinform. | 5 |
| 2008 | Assessment of disease named entity recognition on a corpus of annotated sentencesabstractBACKGROUND: In recent years, the recognition of semantic types from the biomedical scientific literature has been focused on named entities like protein and gene names (PGNs) and gene ontology terms (GO terms). Other semantic types like diseases have not received the same level of attention. Different solutions have been proposed to identify disease named entities in the scientific literature. While matching the terminology with language patterns suffers from low recall (e.g., Whatizit) other solutions make use of morpho-syntactic features to better cover the full scope of terminological variability (e.g., MetaMap). Currently, MetaMap that is provided from the National Library of Medicine (NLM) is the state of the art solution for the annotation of concepts from UMLS (Unified Medical Language System) in the literature. Nonetheless, its performance has not yet been assessed on an annotated corpus. In addition, little effort has been invested so far to generate an annotated dataset that links disease entities in text to disease entries in a database, thesaurus or ontology and that could serve as a gold standard to benchmark text mining solutions. RESULTS: As part of our research work, we have taken a corpus that has been delivered in the past for the identification of associations of genes to diseases based on the UMLS Metathesaurus and we have reprocessed and re-annotated the corpus. We have gathered annotations for disease entities from two curators, analyzed their disagreement (0.51 in the kappa-statistic) and composed a single annotated corpus for public use. Thereafter, three solutions for disease named entity recognition including MetaMap have been applied to the corpus to automatically annotate it with UMLS Metathesaurus concepts. The resulting annotations have been benchmarked to compare their performance. CONCLUSIONS: The annotated corpus is publicly available at ftp://ftp.ebi.ac.uk/pub/software/textmining/corpora/diseases and can serve as a benchmark to other systems. In addition, we found that dictionary look-up already provides competitive results indicating that the use of disease terminology is highly standardized throughout the terminologies and the literature. MetaMap generates precise results at the expense of insufficient recall while our statistical method obtains better recall at a lower precision rate. Even better results in terms of precision are achieved by combining at least two of the three methods leading, but this approach again lowers recall. Altogether, our analysis gives a better understanding of the complexity of disease annotations in the literature. MetaMap and the dictionary based approach are available through the Whatizit web service infrastructure (Rebholz-Schuhmann D, Arregui M, Gaudan S, Kirsch H, Jimeno A: Text processing through Web services: Calling Whatizit. Bioinformatics 2008, 24:296-298). Antonio Jimeno-Yepes, Ernesto Jiménez-Ruiz, Vivian Lee, Sylvain Gaudan, Rafael Berlanga Llavori, Dietrich Rebholz-Schuhmann |
BMC Bioinform. | 1 |
| 2005 | Data-poor categorization and passage retrieval for Gene Ontology Annotation in Swiss-ProtabstractBACKGROUND: In the context of the BioCreative competition, where training data were very sparse, we investigated two complementary tasks: 1) given a Swiss-Prot triplet, containing a protein, a GO (Gene Ontology) term and a relevant article, extraction of a short passage that justifies the GO category assignment; 2) given a Swiss-Prot pair, containing a protein and a relevant article, automatic assignment of a set of categories. METHODS: Sentence is the basic retrieval unit. Our classifier computes a distance between each sentence and the GO category provided with the Swiss-Prot entry. The Text Categorizer computes a distance between each GO term and the text of the article. Evaluations are reported both based on annotator judgements as established by the competition and based on mean average precision measures computed using a curated sample of Swiss-Prot. RESULTS: Our system achieved the best recall and precision combination both for passage retrieval and text categorization as evaluated by official evaluators. However, text categorization results were far below those in other data-poor text categorization experiments The top proposed term is relevant in less that 20% of cases, while categorization with other biomedical controlled vocabulary, such as the Medical Subject Headings, we achieved more than 90% precision. We also observe that the scoring methods used in our experiments, based on the retrieval status value of our engines, exhibits effective confidence estimation capabilities. CONCLUSION: From a comparative perspective, the combination of retrieval and natural language processing methods we designed, achieved very competitive performances. Largely data-independent, our systems were no less effective that data-intensive approaches. These results suggests that the overall strategy could benefit a large class of information extraction tasks, especially when training data are missing. However, from a user perspective, results were disappointing. Further investigations are needed to design applicable end-user text mining tools for biologists. Frédéric Ehrler, Antoine Geissbühler, Antonio Jimeno-Yepes, Patrick Ruch |
BMC Bioinform. | 3 |