James G. Mork

dblp:78/3928 · DBLP profile ↗
← Back
35ranked-venue papers
3as first author
5since 2021 · last 2023
0000-0002-0346-1350ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 34 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2023 The National Library of Medicine indexer assignment dataset: A new large-scale dataset for reviewer assignment research
abstract
MEDLINE is the National Library of Medicine's (NLM) journal citation database. It contains over 28 million references to biomedical and life science journal articles, and a key feature of the database is that all articles are indexed with NLM Medical Subject Headings (MeSH). The library employs a team of MeSH indexers, and in recent years they have been asked to index close to 1 million articles per year in order to keep MEDLINE up to date. An important part of the MEDLINE indexing process is the assignment of articles to indexers. High quality and timely indexing is only possible when articles are assigned to indexers with suitable expertise. This paper introduces the NLM indexer assignment dataset: a large dataset of 4.2 million indexer article assignments for articles indexed between 2011 and 2019. The dataset is shown to be a valuable testbed for expert matching and assignment algorithms, and indexer article assignment is also found to be useful domain-adaptive pre-training for the closely related task of reviewer assignment.
Alastair R. Rae, James G. Mork, Dina Demner-Fushman
J. Assoc. Inf. Sci. Technol.2
2021 Using Deep Neural Network Binary Classifiers to Infer MeSH Publication Types
Victor Cid, James G. Mork
AMIA2
2021 Hybrid Ensemble-Rule Algorithm for Improved MEDLINE® Sentence Boundary Detection
Daniel X. Le, James G. Mork, Sameer Antani
AMIA2
2021 Identifying Criteria for Antonym Generation from Collocates in Corpora
Chris J. Lu, Amanda Payne, James G. Mork
AMIA3
2021 Proper filter usage to retrieve multiwords from the MEDLINE n-gram set: Reply to the Turki et al commentary "Enhancing filter-based parenthetic abbreviation extraction methods"
Chris J. Lu, Amanda Payne, James G. Mork
J. Am. Medical Informatics Assoc.3
2020 Enhanced Features in the SPECIALIST Lexicon - Antonyms
Chris J. Lu, Amanda Payne, James G. Mork
AMIA3
2020 Automatic MeSH Indexing: Revisiting the Subheading Attachment Problem
Alastair R. Rae, David O. Pritchard, James G. Mork, Dina Demner-Fushman
AMIA3
2020 The Unified Medical Language System SPECIALIST Lexicon and Lexical Tools: Development and applications
abstract
Natural language processing (NLP) plays a vital role in modern medical informatics. It converts narrative text or unstructured data into knowledge by analyzing and extracting concepts. A comprehensive lexical system is the foundation to the success of NLP applications and an essential component at the beginning of the NLP pipeline. The SPECIALIST Lexicon and Lexical Tools, distributed by the National Library of Medicine as one of the Unified Medical Language System Knowledge Sources, provides an underlying resource for many NLP applications. This article reports recent developments of 3 key components in the Lexicon. The core NLP operation of Unified Medical Language System concept mapping is used to illustrate the importance of these developments. Our objective is to provide generic, broad coverage and a robust lexical system for NLP applications. A novel multiword approach and other planned developments are proposed.
Chris J. Lu, Amanda Payne, James G. Mork
J. Am. Medical Informatics Assoc.3
2019 A High Recall Classifier for Selecting Articles for MEDLINE Indexing
Alastair R. Rae, Max E. Savery, James G. Mork, Dina Demner-Fushman
AMIA3
2019 Evaluation of System for Selective Indexing Classification
Max E. Savery, Melanie Huston, James G. Mork, Olga Printseva, Susan Schmidt, Alastair R. Rae, Dina Demner-Fushman
AMIA3
2018 Finding medication doses in the literature
Dina Demner-Fushman, James G. Mork, Willie J. Rogers, Sonya E. Shooshan, Laritza Rodriguez, Alan R. Aronson
AMIA2
2018 Increasing UMLS Coverage and Reducing Ambiguity via Automated Creation of Synonymous Terms
François-Michel Lang, James G. Mork, Dina Demner-Fushman, Alan R. Aronson
AMIA2
2016 Resolving Hierarchical Ambiguity in Indexing Recommendations
James G. Mork, Dina Demner-Fushman
AMIA1
2015 Extracting Characteristics of the Study Subjects from Full-Text Articles
Dina Demner-Fushman, James G. Mork
AMIA2
2015 Feature engineering for MEDLINE citation categorization with MeSH
abstract
BACKGROUND: Research in biomedical text categorization has mostly used the bag-of-words representation. Other more sophisticated representations of text based on syntactic, semantic and argumentative properties have been less studied. In this paper, we evaluate the impact of different text representations of biomedical texts as features for reproducing the MeSH annotations of some of the most frequent MeSH headings. In addition to unigrams and bigrams, these features include noun phrases, citation meta-data, citation structure, and semantic annotation of the citations. RESULTS: Traditional features like unigrams and bigrams exhibit strong performance compared to other feature sets. Little or no improvement is obtained when using meta-data or citation structure. Noun phrases are too sparse and thus have lower performance compared to more traditional features. Conceptual annotation of the texts by MetaMap shows similar performance compared to unigrams, but adding concepts from the UMLS taxonomy does not improve the performance of using only mapped concepts. The combination of all the features performs largely better than any individual feature set considered. In addition, this combination improves the performance of a state-of-the-art MeSH indexer. Concerning the machine learning algorithms, we find that those that are more resilient to class imbalance largely obtain better performance. CONCLUSIONS: We conclude that even though traditional features such as unigrams and bigrams have strong performance compared to other features, it is possible to combine them to effectively improve the performance of the bag-of-words representation. We have also found that the combination of the learning algorithm and feature sets has an influence in the overall performance of the system. Moreover, using learning algorithms resilient to class imbalance largely improves performance. However, when using a large set of features, consideration needs to be taken with algorithms due to the risk of over-fitting. Specific combinations of learning algorithms and features for individual MeSH headings could further increase the performance of an indexing system.
Antonio Jimeno-Yepes, Laura Plaza, Jorge Carrillo de Albornoz, James G. Mork, Alan R. Aronson
BMC Bioinform.4
2014 Vocabulary Density Method for Customized Indexing of MEDLINE Journals
James G. Mork, Dina Demner-Fushman, Susan Schmidt, Alan R. Aronson
AMIA1
2013 Mining MEDLINE for problems associated with vitamin D
Dina Demner-Fushman, James G. Mork, Alan R. Aronson
AMIA2
2013 Comparison and combination of several MeSH indexing approaches
Antonio Jimeno-Yepes, James G. Mork, Dina Demner-Fushman, Alan R. Aronson
AMIA2
2013 MeSH indexing based on automatically generated summaries
abstract
BACKGROUND: MEDLINE citations are manually indexed at the U.S. National Library of Medicine (NLM) using as reference the Medical Subject Headings (MeSH) controlled vocabulary. For this task, the human indexers read the full text of the article. Due to the growth of MEDLINE, the NLM Indexing Initiative explores indexing methodologies that can support the task of the indexers. Medical Text Indexer (MTI) is a tool developed by the NLM Indexing Initiative to provide MeSH indexing recommendations to indexers. Currently, the input to MTI is MEDLINE citations, title and abstract only. Previous work has shown that using full text as input to MTI increases recall, but decreases precision sharply. We propose using summaries generated automatically from the full text for the input to MTI to use in the task of suggesting MeSH headings to indexers. Summaries distill the most salient information from the full text, which might increase the coverage of automatic indexing approaches based on MEDLINE. We hypothesize that if the results were good enough, manual indexers could possibly use automatic summaries instead of the full texts, along with the recommendations of MTI, to speed up the process while maintaining high quality of indexing results. RESULTS: We have generated summaries of different lengths using two different summarizers, and evaluated the MTI indexing on the summaries using different algorithms: MTI, individual MTI components, and machine learning. The results are compared to those of full text articles and MEDLINE citations. Our results show that automatically generated summaries achieve similar recall but higher precision compared to full text articles. Compared to MEDLINE citations, summaries achieve higher recall but lower precision. CONCLUSIONS: Our results show that automatic summaries produce better indexing than full text articles. Summaries produce similar recall to full text but much better precision, which seems to indicate that automatic summaries can efficiently capture the most important contents within the original articles. The combination of MEDLINE citations and automatically generated summaries could improve the recommendations suggested by MTI. On the other hand, indexing performance might be dependent on the MeSH heading being indexed. Summarization techniques could thus be considered as a feature selection algorithm that might have to be tuned individually for each MeSH heading.
Antonio Jimeno-Yepes, Laura Plaza, James G. Mork, Alan R. Aronson, Alberto Díaz 0001
BMC Bioinform.3
2013 GeneRIF indexing: sentence selection based on machine learning
abstract
BACKGROUND: A Gene Reference Into Function (GeneRIF) describes novel functionality of genes. GeneRIFs are available from the National Center for Biotechnology Information (NCBI) Gene database. GeneRIF indexing is performed manually, and the intention of our work is to provide methods to support creating the GeneRIF entries. The creation of GeneRIF entries involves the identification of the genes mentioned in MEDLINE®; citations and the sentences describing a novel function. RESULTS: We have compared several learning algorithms and several features extracted or derived from MEDLINE sentences to determine if a sentence should be selected for GeneRIF indexing. Features are derived from the sentences or using mechanisms to augment the information provided by them: assigning a discourse label using a previously trained model, for example. We show that machine learning approaches with specific feature combinations achieve results close to one of the annotators. We have evaluated different feature sets and learning algorithms. In particular, Naïve Bayes achieves better performance with a selection of features similar to one used in related work, which considers the location of the sentence, the discourse of the sentence and the functional terminology in it. CONCLUSIONS: The current performance is at a level similar to human annotation and it shows that machine learning can be used to automate the task of sentence selection for GeneRIF annotation. The current experiments are limited to the human species. We would like to see how the methodology can be extended to other species, specifically the normalization of gene mentions in other species.
Antonio Jimeno-Yepes, J. Caitlin Sticco, James G. Mork, Alan R. Aronson
BMC Bioinform.3
2012 Structured Abstracts in MEDLINE: Implementation Based on a Retrospective Cohort Study
Anna Ripple, Lou Knecht, John Rozier, James G. Mork
AMIA4
2010 Extracting Rx information from clinical narrative
abstract
OBJECTIVE: The authors used the i2b2 Medication Extraction Challenge to evaluate their entity extraction methods, contribute to the generation of a publicly available collection of annotated clinical notes, and start developing methods for ontology-based reasoning using structured information generated from the unstructured clinical narrative. DESIGN: Extraction of salient features of medication orders from the text of de-identified hospital discharge summaries was addressed with a knowledge-based approach using simple rules and lookup lists. The entity recognition tool, MetaMap, was combined with dose, frequency, and duration modules specifically developed for the Challenge as well as a prototype module for reason identification. MEASUREMENTS: Evaluation metrics and corresponding results were provided by the Challenge organizers. RESULTS: The results indicate that robust rule-based tools achieve satisfactory results in extraction of simple elements of medication orders, but more sophisticated methods are needed for identification of reasons for the orders and durations. LIMITATIONS: Owing to the time constraints and nature of the Challenge, some obvious follow-on analysis has not been completed yet. CONCLUSIONS: The authors plan to integrate the new modules with MetaMap to enhance its accuracy. This integration effort will provide guidance in retargeting existing tools for better processing of clinical text.
James G. Mork, Olivier Bodenreider, Dina Demner-Fushman, Rezarta Islamaj Dogan, François-Michel Lang, Zhiyong Lu, Aurélie Névéol, Lee B. Peters, Sonya E. Shooshan, Alan R. Aronson
J. Am. Medical Informatics Assoc.1
2010 UMLS content views appropriate for NLP processing of the biomedical literature vs. clinical text
Dina Demner-Fushman, James G. Mork, Sonya E. Shooshan, Alan R. Aronson
J. Biomed. Informatics2
2009 Comment on 'MeSH-up: effective MeSH text classification for improved document retrieval'
abstract
Abstract Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Aurélie Névéol, James G. Mork, Alan R. Aronson
Bioinform.2
2009 A recent advance in the automatic indexing of the biomedical literature
Aurélie Névéol, Sonya E. Shooshan, Susanne M. Humphrey, James G. Mork, Alan R. Aronson
J. Biomed. Informatics4
2008 Methodology for Creating UMLS Content Views Appropriate for Biomedical Natural Language Processing
Alan R. Aronson, James G. Mork, Aurélie Névéol, Sonya E. Shooshan, Dina Demner-Fushman
AMIA2
2007 Fine-Grained Indexing of the Biomedical Literature: MeSH Subheading Attachment for a MEDLINE Indexing Tool
Aurélie Névéol, Sonya E. Shooshan, James G. Mork, Alan R. Aronson
AMIA3
2005 Evaluation of French and English MeSH Indexing Systems with a Parallel Corpus
Aurélie Névéol, James G. Mork, Alan R. Aronson, Stéfan Jacques Darmoni
AMIA2
2003 Gene Indexing: Characterization and Analysis of NLM's GeneRIFs
Joyce A. Mitchell, Alan R. Aronson, James G. Mork, Lillian C. Folk, Susanne M. Humphrey, Janice M. Ward
AMIA3
2002 Finding UMLS Metathesaurus concepts in MEDLINE
Suresh Srinivasan, Thomas C. Rindflesch, William T. Hole, Alan R. Aronson, James G. Mork
AMIA5
2001 Developing a test collection for biomedical word sense disambiguation
Marc Weeber, James G. Mork, Alan R. Aronson
AMIA2
2000 The NLM Indexing Initiative
Alan R. Aronson, Olivier Bodenreider, Florence Chang, Susanne M. Humphrey, James G. Mork, Stuart J. Nelson, Thomas C. Rindflesch, W. John Wilbur
AMIA5
2000 Text-based discovery in biomedicine: the architecture of the DAD-system
Marc Weeber, Henny Klein, Alan R. Aronson, James G. Mork, Lolkje T. W. de Jong-van den Berg, Rein Vos
AMIA4
1999 Automated Assignment of Medical Subject Headings
Stuart J. Nelson, Alan R. Aronson, Tamas E. Doszkocs, W. John Wilbur, Olivier Bodenreider, Florence Chang, James G. Mork, Alexa T. McCray
AMIA7
1999 Analysis of biomedical text for chemical names: a comparison of three methods
W. John Wilbur, George F. Hazard Jr., Guy Divita, James G. Mork, Alan R. Aronson, Allen C. Browne
AMIA4