VLDB 2026 Research / reviewers in the wild / expert
Frank J. Manion
dblp:53/2658
· DBLP profile ↗
24ranked-venue papers
2as first author
4since 2021 · last 2024
0000-0003-1030-6348ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 24 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | AutoCriteria: a generalizable clinical trial eligibility criteria extraction system powered by large language modelsabstractOBJECTIVES: We aim to build a generalizable information extraction system leveraging large language models to extract granular eligibility criteria information for diverse diseases from free text clinical trial protocol documents. We investigate the model's capability to extract criteria entities along with contextual attributes including values, temporality, and modifiers and present the strengths and limitations of this system. MATERIALS AND METHODS: The clinical trial data were acquired from https://ClinicalTrials.gov/. We developed a system, AutoCriteria, which comprises the following modules: preprocessing, knowledge ingestion, prompt modeling based on GPT, postprocessing, and interim evaluation. The final system evaluation was performed, both quantitatively and qualitatively, on 180 manually annotated trials encompassing 9 diseases. RESULTS: AutoCriteria achieves an overall F1 score of 89.42 across all 9 diseases in extracting the criteria entities, with the highest being 95.44 for nonalcoholic steatohepatitis and the lowest of 84.10 for breast cancer. Its overall accuracy is 78.95% in identifying all contextual information across all diseases. Our thematic analysis indicated accurate logic interpretation of criteria as one of the strengths and overlooking/neglecting the main criteria as one of the weaknesses of AutoCriteria. DISCUSSION: AutoCriteria demonstrates strong potential to extract granular eligibility criteria information from trial documents without requiring manual annotations. The prompts developed for AutoCriteria generalize well across different disease areas. Our evaluation suggests that the system handles complex scenarios including multiple arm conditions and logics. CONCLUSION: AutoCriteria currently encompasses a diverse range of diseases and holds potential to extend to more in the future. This signifies a generalizable and scalable solution, poised to address the complexities of clinical trial application in real-world settings. Surabhi Datta, Kyeryoung Lee, Hunki Paek, Frank J. Manion, Nneka Ofoegbu, Jingcheng Du, Liang-Chin Huang, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 4 |
| 2023 | Machine learning-based donor permission extraction from informed consent documentsabstractBACKGROUND: With more clinical trials are offering optional participation in the collection of bio-specimens for biobanking comes the increasing complexity of requirements of informed consent forms. The aim of this study is to develop an automatic natural language processing (NLP) tool to annotate informed consent documents to promote biorepository data regulation, sharing, and decision support. We collected informed consent documents from several publicly available sources, then manually annotated them, covering sentences containing permission information about the sharing of either bio-specimens or donor data, or conducting genetic research or future research using bio-specimens or donor data. RESULTS: We evaluated a variety of machine learning algorithms including random forest (RF) and support vector machine (SVM) for the automatic identification of these sentences. 120 informed consent documents containing 29,204 sentences were annotated, of which 1250 sentences (4.28%) provide answers to a permission question. A support vector machine (SVM) model achieved a F-1 score of 0.95 on classifying the sentences when using a gold standard, which is a prefiltered corpus containing all relevant sentences. CONCLUSIONS: This study provides the feasibility of using machine learning tools to classify permission-related sentences in informed consent documents. Madhuri Sankaranarayanapillai, Jingcheng Du, Yang Xiang 0003, Frank J. Manion, Marcelline R. Harris, Cooper Stansbury, Huy Anh Pham, Cui Tao |
BMC Bioinform. | 5 |
| 2021 | Expressing and Executing Informed Consent Permissions using SWRL: The All of Us Use Case
Muhammad Amith, Marcelline R. Harris, Cooper Stansbury, Kathleen Ford, Frank J. Manion, Cui Tao |
AMIA | 5 |
| 2021 | COVID-19 SignSym: a fast adaptation of a general clinical NLP tool to identify and normalize COVID-19 signs and symptoms to OMOP common data modelabstractThe COVID-19 pandemic swept across the world rapidly, infecting millions of people. An efficient tool that can accurately recognize important clinical concepts of COVID-19 from free text in electronic health records (EHRs) will be valuable to accelerate COVID-19 clinical research. To this end, this study aims at adapting the existing CLAMP natural language processing tool to quickly build COVID-19 SignSym, which can extract COVID-19 signs/symptoms and their 8 attributes (body location, severity, temporal expression, subject, condition, uncertainty, negation, and course) from clinical text. The extracted information is also mapped to standard concepts in the Observational Medical Outcomes Partnership common data model. A hybrid approach of combining deep learning-based models, curated lexicons, and pattern-based rules was applied to quickly build the COVID-19 SignSym from CLAMP, with optimized performance. Our extensive evaluation using 3 external sites with clinical notes of COVID-19 patients, as well as the online medical dialogues of COVID-19, shows COVID-19 SignSym can achieve high performance across data sources. The workflow used for this study can be generalized to other use cases, where existing clinical natural language processing tools need to be customized for specific information needs within a short time. COVID-19 SignSym is freely accessible to the research community as a downloadable package (https://clamp.uth.edu/covid/nlp.php) and has been used by 16 healthcare organizations to support clinical research of COVID-19. Noor Abu-El-Rub, Josh Gray, Huy Anh Pham, Yujia Zhou 0003, Frank J. Manion, Xing Song, Hua Xu 0001, Masoud Rouhizadeh, Yaoyun Zhang |
J. Am. Medical Informatics Assoc. | 6 |
| 2019 | A Scoping Review to Identify Biobank Classification and Permissions to Share Metadata
Marcelline R. Harris, Frank J. Manion, Madhuri Sankaranarayanapillai, Hsing-yi Song, Marisa Conte, Cui Tao |
AMIA | 2 |
| 2019 | Refactoring and Expanding the Informed Consent Ontology (ICO)
Jonathan Vajda, J. Neil Otte, Cooper Stansbury, Marcelline R. Harris, Frank J. Manion, Cui Tao |
AMIA | 5 |
| 2019 | Expanding the Representation of Permissions and Deontic Roles in the Informed Consent Ontology (ICO)
Jonathan Vajda, J. Neil Otte, Cooper Stansbury, Marcelline R. Harris, Elizabeth Umberfield, Frank J. Manion, Cui Tao |
AMIA | 6 |
| 2018 | Development and Validation of an Ontological Representation of the US Common Rule
Frank J. Manion, Marcelline R. Harris, Muhamamd F. Amith, Cui Tao |
AMIA | 1 |
| 2018 | OntoKeeper: Semiotic-driven Ontology Evaluation Tool For Biomedical Ontologists
Muhammad Amith, Frank J. Manion, Chen Liang 0005, Marcelline R. Harris, Dennis Wang, Yongqun He, Cui Tao |
BIBM | 2 |
| 2017 | From Scoping Review to Metadata: An Evidence-Based Approach to Identifying Informed Consent Metadata for Biorepositories
Marcelline R. Harris, Frank J. Manion, Hsing-yi Song, Yongqun He, Muhamamd F. Amith, Cui Tao |
AMIA | 2 |
| 2017 | Evaluation of the Surveillance, Epidemiology, and End Results Data Management System (SEER*DMS) to support the efforts in enhancing U.S cancer surveillance
Marina Z. Matatova, Paul Fearn, Eric B. Durbin, Gary Levin, Frank J. Manion, Patrick Mergler |
AMIA | 5 |
| 2016 | A Research Capability Framework
Airong Luo, Marcelline R. Harris, Frank J. Manion, Barbara Mirel |
AMIA | 3 |
| 2014 | Analysis of Content Coverage for Informed Consent Concepts
Frank J. Manion, Elizabeth Eisenhauer, Alla Karnovsky, Yongqun He, Marcelline R. Harris |
AMIA | 1 |
| 2012 | Extending OpenClinica using Web 2.0 Technologies
Anjali Bhandari, John Harju, Frank J. Manion, Becky Marshall, Rebecca M. Minter |
AMIA | 4 |
| 2012 | Hedging their Mets: The Use of Uncertainty Terms in Clinical Documents and its Potential Implications when Sharing the Documents with Patients
David A. Hanauer, Yang Liu 0019, Qiaozhu Mei, Frank J. Manion, Ulysses J. Balis, Kai Zheng 0002 |
AMIA | 4 |
| 2012 | Development of an Informed Consent Ontology to Support Biobanking
Alla Karnovsky, Frank J. Manion, Terry E. Weymouth, V. Glenn Tarcea, Lisa Powell, Blake Roessler, Nicholas Steneck |
AMIA | 2 |
| 2012 | tranSMART Supports a Post-GWAS Data Coordinating Center
Dan Stuart, John Harju, David A. Hanauer, Dina Aronzon, Raveen Sharma, Frank J. Manion, Haiping Xia, Carolyn Hutter, Stephen Gruber |
AMIA | 8 |
| 2009 | FuGEFlow: data model and markup language for flow cytometryabstractBACKGROUND: Flow cytometry technology is widely used in both health care and research. The rapid expansion of flow cytometry applications has outpaced the development of data storage and analysis tools. Collaborative efforts being taken to eliminate this gap include building common vocabularies and ontologies, designing generic data models, and defining data exchange formats. The Minimum Information about a Flow Cytometry Experiment (MIFlowCyt) standard was recently adopted by the International Society for Advancement of Cytometry. This standard guides researchers on the information that should be included in peer reviewed publications, but it is insufficient for data exchange and integration between computational systems. The Functional Genomics Experiment (FuGE) formalizes common aspects of comprehensive and high throughput experiments across different biological technologies. We have extended FuGE object model to accommodate flow cytometry data and metadata. METHODS: We used the MagicDraw modelling tool to design a UML model (Flow-OM) according to the FuGE extension guidelines and the AndroMDA toolkit to transform the model to a markup language (Flow-ML). We mapped each MIFlowCyt term to either an existing FuGE class or to a new FuGEFlow class. The development environment was validated by comparing the official FuGE XSD to the schema we generated from the FuGE object model using our configuration. After the Flow-OM model was completed, the final version of the Flow-ML was generated and validated against an example MIFlowCyt compliant experiment description. RESULTS: The extension of FuGE for flow cytometry has resulted in a generic FuGE-compliant data model (FuGEFlow), which accommodates and links together all information required by MIFlowCyt. The FuGEFlow model can be used to build software and databases using FuGE software toolkits to facilitate automated exchange and manipulation of potentially large flow cytometry experimental data sets. Additional project documentation, including reusable design patterns and a guide for setting up a development environment, was contributed back to the FuGE project. CONCLUSION: We have shown that an extension of FuGE can be used to transform minimum information requirements in natural language to markup language in XML. Extending FuGE required significant effort, but in our experiences the benefits outweighed the costs. The FuGEFlow is expected to play a central role in describing flow cytometry experiments and ultimately facilitating data exchange including public flow cytometry repositories currently under development. Olga Tchuvatkina, Josef Spidlen, Peter Wilkinson, Maura Gasparetto, Andrew R. Jones, Frank J. Manion, Richard H. Scheuermann, Rafick-Pierre Sekaly, Ryan Remy Brinkman |
BMC Bioinform. | 7 |
| 2009 | Integration of prostate cancer clinical data using an ontology
Hua Min, Frank J. Manion, Elizabeth Goralczyk, Yu-Ning Wong, Eric A. Ross, J. Robert Beck |
J. Biomed. Informatics | 2 |
| 2006 | WaveRead: Automatic measurement of relative gene expression levels from microarrays using wavelet analysis
Ghislain Bidaut, Frank J. Manion, Christophe Garcia, Michael F. Ochs |
J. Biomed. Informatics | 2 |
| 2004 | FGDP: functional genomics data pipeline for automated, multiple microarray data analysesabstractUNLABELLED: Gene expression microarrays and oligonucleotide GeneChips have provided biologists with a means of measuring, in a single experiment, the expression levels of entire genomes under a variety of conditions. As with any nascent field, there is no single accepted method for analyzing the new data types, with new methods appearing monthly. Investigators using the new technology must constantly seek access to the latest tools and explore their data in multiple ways. The functional genomics data pipeline provides an integrated, extendable analysis environment permitting multiple, simultaneous analyses to be automatically performed and provides a web server and interface for presenting results. AVAILABILITY: Source code and executables are available under the GNU public license at http://bioinformatics.fccc.edu/ Jeffrey D. Grant, Luke A. Somers, Frank J. Manion, Ghislain Bidaut, Michael F. Ochs |
Bioinform. | 4 |
| 2003 | ASAP: automated sequence annotation pipeline for web-based updating of sequence information with a local dynamic databaseabstractAbstract Summary: The automated sequence annotation pipeline (ASAP) is designed to ease routine investigation of new functional annotations on unknown sequences, such as expressed sequence tags (ESTs), through querying of web-accessible resources and maintenance of a local database. The system allows easy use of the output from one search as the input for a new search, as well as the filtering of results. The database is used to store formats and parameters and information for parsing data from web sites. The database permits easy updating of format information should a site modify the format of a query or of a returned web page. Availability: Source code is available under the GNU public license at http://bioinformatics.fccc.edu/ Contact: [email protected][email protected]@[email protected][email protected] * To whom correspondence should be addressed. Andrew V. Kossenkov, Frank J. Manion, Eugene V. Korotkov, Thomas D. Moloshok, Michael F. Ochs |
Bioinform. | 2 |
| 2002 | BeoBLAST: distributed BLAST and PSI-BLAST on a Beowulf clusterabstractUNLABELLED: BeoBLAST is an integrated software package that handles user requests and distributes BLAST and PSI-BLAST searches to nodes of a Beowulf cluster, thus providing a simple way to implement a scalable BLAST system on top of relatively inexpensive computer clusters. Additionally, BeoBLAST offers a number of novel search features through its web interface, including the ability to perform simultaneous searches of multiple databases with multiple queries, and the ability to start a search using the PSSM generated from a previous PSI-BLAST search on a different database. The underlying system can also handle automated querying for high throughput work. AVAILABILITY: Source code is available under the GNU public license at http://bioinformatics.fccc.edu/ Jeffrey D. Grant, Roland L. Dunbrack Jr., Frank J. Manion, Michael F. Ochs |
Bioinform. | 3 |
| 2002 | Application of Bayesian Decomposition for analysing microarray dataabstractMOTIVATION: Microarray and gene chip technology provide high throughput tools for measuring gene expression levels in a variety of circumstances, including cellular response to drug treatment, cellular growth and development, tumorigenesis, among many other processes. In order to interpret the large data sets generated in experiments, data analysis techniques that consider biological knowledge during analysis will be extremely useful. We present here results showing the application of such a tool to expression data from yeast cell cycle experiments. RESULTS: Originally developed for spectroscopic analysis, Bayesian Decomposition (BD) includes two features which make it useful for microarray data analysis: the ability to assign genes to multiple coexpression groups and the ability to encode biological knowledge into the system. Here we demonstrate the ability of the algorithm to provide insight into the yeast cell cycle, including identification of five temporal patterns tied to cell cycle phases as well as the identification of a pattern tied to an approximately 40 min cell cycle oscillator. The genes are simultaneously assigned to the patterns, including partial assignment to multiple patterns when this is required to explain the expression profile. AVAILABILITY: The application is available free to academic users under a material transfer agreement. Go to http://bioinformatics.fccc.edu/ for more details. Thomas D. Moloshok, R. R. Klevecz, Jeffrey D. Grant, Frank J. Manion, W. F. Speier IV, Michael F. Ochs |
Bioinform. | 4 |