EDBT 2026 Demo / reviewers in the wild / expert
Scott L. DuVall
dblp:11/4887
· DBLP profile ↗
53ranked-venue papers
5as first author
10since 2021 · last 2023
0000-0002-4898-3865ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 51 · 5 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Blockchain-enabled immutable, distributed, and highly available clinical research activity logging system for federated COVID-19 data analysis from multiple institutionsabstractOBJECTIVE: We aimed to develop a distributed, immutable, and highly available cross-cloud blockchain system to facilitate federated data analysis activities among multiple institutions. MATERIALS AND METHODS: We preprocessed 9166 COVID-19 Structured Query Language (SQL) code, summary statistics, and user activity logs, from the GitHub repository of the Reliable Response Data Discovery for COVID-19 (R2D2) Consortium. The repository collected local summary statistics from participating institutions and aggregated the global result to a COVID-19-related clinical query, previously posted by clinicians on a website. We developed both on-chain and off-chain components to store/query these activity logs and their associated queries/results on a blockchain for immutability, transparency, and high availability of research communication. We measured run-time efficiency of contract deployment, network transactions, and confirmed the accuracy of recorded logs compared to a centralized baseline solution. RESULTS: The smart contract deployment took 4.5 s on an average. The time to record an activity log on blockchain was slightly over 2 s, versus 5-9 s for baseline. For querying, each query took on an average less than 0.4 s on blockchain, versus around 2.1 s for baseline. DISCUSSION: The low deployment, recording, and querying times confirm the feasibility of our cross-cloud, blockchain-based federated data analysis system. We have yet to evaluate the system on a larger network with multiple nodes per cloud, to consider how to accommodate a surge in activities, and to investigate methods to lower querying time as the blockchain grows. CONCLUSION: Blockchain technology can be used to support federated data analysis among multiple institutions. Tsung-Ting Kuo, Anh Pham, Maxim E. Edelson, Jihoon Kim 0001, Yash Gupta, Lucila Ohno-Machado, David M. Anderson, Chandrasekar Balacha, Tyler Bath, Sally L. Baxter, Andrea Becker-Pennrich, Douglas S. Bell, Elmer V. Bernstam, Ngan Chau, Michele E. Day, Jason N. Doctor, Scott L. DuVall, Robert El-Kareh, Renato Florian, Robert W. Follett, Benjamin P. Geisler, Alessandro Ghigi, Assaf Gottlieb, Christian Hinske, Zhaoxian Hu, Diana Ir, Xiaoqian Jiang, Katherine K. Kim, Tara K. Knight, Jejo Koola, Ulrich Mansmann, Michael E. Matheny, Daniella Meeker, Zongyang Mou, Larissa Neumann, Nghia H. Nguyen, Nicholas R. Anderson 0001, Eunice Park, Paulina Paul, Mark J. Pletcher, Kai W. Post, Clemens Rieder, Clemens Scherer, Lisa M. Schilling, Andrey Soares, Spencer L. SooHoo, Ekin Soysal, Steven Covington, Brian Tep, Brian Toy, Baocheng Wang, Zhen R. Wu, Hua Xu 0001, Yong K. Choi, Kai Zheng 0002, Yujia Zhou 0003, Rachel A Zucker |
J. Am. Medical Informatics Assoc. | 18 |
| 2023 | Reproducible variability: assessing investigator discordance across 9 research teams attempting to reproduce the same observational studyabstractOBJECTIVE: Observational studies can impact patient care but must be robust and reproducible. Nonreproducibility is primarily caused by unclear reporting of design choices and analytic procedures. This study aimed to: (1) assess how the study logic described in an observational study could be interpreted by independent researchers and (2) quantify the impact of interpretations' variability on patient characteristics. MATERIALS AND METHODS: Nine teams of highly qualified researchers reproduced a cohort from a study by Albogami et al. The teams were provided the clinical codes and access to the tools to create cohort definitions such that the only variable part was their logic choices. We executed teams' cohort definitions against the database and compared the number of subjects, patient overlap, and patient characteristics. RESULTS: On average, the teams' interpretations fully aligned with the master implementation in 4 out of 10 inclusion criteria with at least 4 deviations per team. Cohorts' size varied from one-third of the master cohort size to 10 times the cohort size (2159-63 619 subjects compared to 6196 subjects). Median agreement was 9.4% (interquartile range 15.3-16.2%). The teams' cohorts significantly differed from the master implementation by at least 2 baseline characteristics, and most of the teams differed by at least 5. CONCLUSIONS: Independent research teams attempting to reproduce the study based on its free-text description alone produce different implementations that vary in the population size and composition. Sharing analytical code supported by a common data model and open-source tools allows reproducing a study unambiguously thereby preserving initial design choices. Anna Ostropolets, Yasser Albogami, Mitchell Conover, Juan M. Banda, William A. Baumgartner Jr., Clair Blacketer, Priyamvada Desai, Scott L. DuVall, Stephen P. Fortin, James P. Gilbert, Asieh Golozar, Joshua Ide, Andrew S. Kanter, David M. Kern, Chungsoo Kim, Lana Y. H. Lai, Kristine E. Lynch, Evan P. Minty, Maria Inês Neves, Ding Quan Ng, Tontel Obene, Victor Pera, Nicole Pratt, Gowtham Rao, Nadav Rappoport, Ines Reinecke, Paola Saroufim, Azza Shoaibi, Katherine Simon, Marc A. Suchard, Joel N. Swerdel, Erica A. Voss, James Weaver, Linying Zhang, George Hripcsak, Patrick B. Ryan |
J. Am. Medical Informatics Assoc. | 8 |
| 2023 | A deep learning approach for medication disposition and corresponding attributes extractionabstractOBJECTIVE: This article summarizes our approach to extracting medication and corresponding attributes from clinical notes, which is the focus of track 1 of the 2022 National Natural Language Processing (NLP) Clinical Challenges(n2c2) shared task. METHODS: The dataset was prepared using Contextualized Medication Event Dataset (CMED), including 500 notes from 296 patients. Our system consisted of three components: medication named entity recognition (NER), event classification (EC), and context classification (CC). These three components were built using transformer models with slightly different architecture and input text engineering. A zero-shot learning solution for CC was also explored. RESULTS: Our best performance systems achieved micro-average F1 scores of 0.973, 0.911, and 0.909 for the NER, EC, and CC, respectively. CONCLUSION: In this study, we implemented a deep learning-based NLP system and demonstrated that our approach of (1) utilizing special tokens helps our model to distinguish multiple medications mentions in the same context; (2) aggregating multiple events of a single medication into multiple labels improves our model's performance. Qiwei Gan, Mengke Hu, Kelly S. Peterson, Hannah Eyre, Patrick R. Alba, Annie E. Bowles, Johnathan C. Stanley, Scott L. DuVall, Jianlin Shi |
J. Biomed. Informatics | 8 |
| 2022 | From Rules to Machine Learning: Upgrading Aging Clinical NLP Systems
Hannah Eyre, Scott L. DuVall, Olga V. Patterson |
AMIA | 2 |
| 2022 | Identifying Menopausal Status with Natural Language Processing
Hannah Eyre, Kristine W. Lynch, Carolyn Gibson, Scott L. DuVall, Olga V. Patterson |
AMIA | 4 |
| 2022 | Manual Chart Review Train and Inter-Annotator Agreement Plan for a COVID-19 Disease Association Study
Brent D. Hill, Kristine W. Lynch, Scott L. DuVall |
AMIA | 3 |
| 2022 | Comprehensive Mapping of Presenting Symptoms in the Emergency Department and Inpatient Setting
Christopher R. Wilson, Annie E. Bowles, Hannah Eyre, Scott L. DuVall, Olga V. Patterson |
AMIA | 4 |
| 2021 | Challenges of comprehensive automatic coding of presenting complaints
Annie E. Bowles, Hannah Eyre, Scott L. DuVall, Olga V. Patterson |
AMIA | 3 |
| 2021 | Launching into clinical space with medspaCy: a new clinical text processing toolkit in Python
Hannah Eyre, Alec B. Chapman, Kelly S. Peterson, Jianlin Shi, Patrick R. Alba, Makoto Jones, Tamara L. Box, Scott L. DuVall, Olga V. Patterson |
AMIA | 8 |
| 2021 | From Emergency Department to Admission: mapping reasons for visit and admit diagnosis using Natural Language Processing
Olga V. Patterson, Hannah Eyre, Kelly S. Peterson, Scott L. DuVall |
AMIA | 4 |
| 2020 | Removing barriers to clinical text processing with MedSpaCy
Hannah Eyre, Olga V. Patterson, Jianlin Shi, Kelly S. Peterson, Alec B. Chapman, Patrick R. Alba, Scott L. DuVall |
AMIA | 7 |
| 2020 | Linking Polyps to Jars: information loss across colonoscopy and pathology reports
Olga V. Patterson, Samir Gupta, Andrew Gawron, Tonya Kaltenbach, Ranier Bustamante, Daniel W. Denhalter, Ashley Earles, Scott L. DuVall |
AMIA | 9 |
| 2020 | COVID-19 TestNorm: A tool to normalize COVID-19 testing names to LOINC codesabstractLarge observational data networks that leverage routine clinical practice data in electronic health records (EHRs) are critical resources for research on coronavirus disease 2019 (COVID-19). Data normalization is a key challenge for the secondary use of EHRs for COVID-19 research across institutions. In this study, we addressed the challenge of automating the normalization of COVID-19 diagnostic tests, which are critical data elements, but for which controlled terminology terms were published after clinical implementation. We developed a simple but effective rule-based tool called COVID-19 TestNorm to automatically normalize local COVID-19 testing names to standard LOINC (Logical Observation Identifiers Names and Codes) codes. COVID-19 TestNorm was developed and evaluated using 568 test names collected from 8 healthcare systems. Our results show that it could achieve an accuracy of 97.4% on an independent test set. COVID-19 TestNorm is available as an open-source package for developers and as an online Web application for end users (https://clamp.uth.edu/covid/loinc.php). We believe that it will be a useful tool to support secondary use of EHRs for research on COVID-19. Jianfu Li, Ekin Soysal, Jiang Bian 0001, Scott L. DuVall, Elizabeth Hanchrow, Kristine E. Lynch, Michael E. Matheny, Karthik Natarajan, Lucila Ohno-Machado, Serguei V. S. Pakhomov, Ruth M. Reeves, Amy M. Sitapati, Swapna Abhyankar, Theresa A. Cullen, Jami Deckard, Xiaoqian Jiang, Robert Murphy, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 5 |
| 2019 | Developing Synthetic VA Healthcare Data in OMOP CDM Model
Jiantao Bian, Hamid Saoudian, Brett R. South, Kristine E. Lynch, Benjamin Viernes, Michael E. Matheny, Scott L. DuVall |
AMIA | 7 |
| 2018 | Ankle Brachial Index Extraction System
Patrick R. Alba, Scott L. DuVall, Daniel Norvell, Kathryn P. Moore, Joseph M. Czerniecki, Olga V. Patterson |
AMIA | 2 |
| 2018 | A User-Friendly Interface for Concept Dictionary Expansion Using Word Embeddings and SNOMED-CT
Alec B. Chapman, Patrick R. Alba, Brian T. Bucher, Scott L. DuVall, Olga V. Patterson |
AMIA | 4 |
| 2018 | The Department of Defense (DoD) and Department of Veterans Affairs (VA) Infrastructure for Clinical Intelligence (DaVINCI)
Scott L. DuVall, Michael E. Matheny, Ildar R. Ibragimov, Trey D. Oats, Jay N. Tucker, Brett R. South, Augie Turano, Hamid Saoudian, Casey Kangas, Keith D. Hofmann, Wendy Funk, Chris Nichols, Albert Bonnema, Louis Ferrucci, Jonathan R. Nebeker |
AMIA | 1 |
| 2018 | Towards intuitive NLP: Interactive system for machine teaching
Olga V. Patterson, Lalindra De Silva, Brad Adams, Ryan Cornia, Thomas Ginter, Susan L. Zickmund, Scott L. DuVall |
AMIA | 7 |
| 2018 | Fast and Accurate Adverse Drug Event labeling without a GPU
Kelly S. Peterson, Alec B. Chapman, Patrick R. Alba, Scott L. DuVall, Olga V. Patterson |
AMIA | 4 |
| 2018 | Flipping the Model for Biomedical Informatics Research
Brett R. South, Kristine W. Lynch, Michael E. Matheny, Julie A. Lynch, Olga Efimova, Catherine Chanfreau-Coffinier, Olga V. Patterson, Benjamin Viernes, Scott L. DuVall |
AMIA | 9 |
| 2017 | Pathology Information Extraction in Bladder Cancer Surveillance
Patrick R. Alba, Scott L. DuVall, Terri Elizabeth Workman, Florian Schroeck, Olga V. Patterson |
AMIA | 2 |
| 2017 | ProjectFlow: Configurable Clinical Trial Management with Enterprise Data Integration and Point-of-Care Study Support
Ryan Cornia, Nilla Majahalme, Danne C. Elbers, Svitlana Dipietro, Brian R. Ivie, Brad Adams, Ramana Seerapu, Valmeek Kudesia, Scott L. DuVall |
AMIA | 9 |
| 2017 | Improving the Quality of Clinical Data Extracted from Text
Daniel W. Denhalter, Olga V. Patterson, Brian D. Robison, Scott L. DuVall |
AMIA | 4 |
| 2017 | For the Common Good: Sharing Data Extracted from Text
Olga V. Patterson, Scott L. DuVall |
AMIA | 2 |
| 2017 | Interactive Visualization and Exploration of Patient Progression in a Hospital Setting
Wathsala Widanagamaachchi, Yarden Livnat, Peer-Timo Bremer, Scott L. DuVall, Valerio Pascucci |
AMIA | 4 |
| 2016 | The Super Annotator: A Method of Semi-Automated Rare Event Identification for Large Clinical Data Sets
Patrick R. Alba, Olga V. Patterson, Benjamin Viernes, Daniel W. Denhalter, Nicole Bailey, Aaron W. C. Kamauu, Scott L. DuVall |
AMIA | 8 |
| 2016 | Large Scale Clinical Text Processing and Process Optimization
Scott L. DuVall, Patrick R. Alba, Olga V. Patterson |
AMIA | 1 |
| 2016 | VIP: A Framework for Mining Clinical Concepts Using Knowledge Author Ontologies and Apache UIMA
Thomas Ginter, Lalindra De Silva, Olga V. Patterson, William Scuba, Wendy W. Chapman, Scott L. DuVall |
AMIA | 6 |
| 2016 | An Introduction to Natural Language Processing Methods in Clinical Research
Olga V. Patterson, Patrick R. Alba, Scott L. DuVall |
AMIA | 3 |
| 2016 | Evaluation of UMLS Term Coverage for Echocardiogram Measures
Olga V. Patterson, Matthew Freiberg, Cynthia Brandt, Scott L. DuVall |
AMIA | 4 |
| 2016 | Automatic Extraction of Maximum and Recommended Drug Dosage Information from DailyMed Database
Lalindra De Silva, Olga V. Patterson, Scott L. DuVall |
AMIA | 3 |
| 2015 | Knowledge Base Acquisition For Rare Concepts Using Manual Bootstrapping
Patrick R. Alba, Scott L. DuVall, Joanne LaFleur, Adam Bress, Olga V. Patterson |
AMIA | 2 |
| 2015 | Transforming the National Department of Veterans Affairs Data Warehouse to the OMOP Common Data Model
Fern FitzHenry, Jesse Brannen, Jason N. Denton, Jonathan R. Nebeker, Scott L. DuVall, Freneka F. Minter, Jeffrey Scehnet, Brian C. Sauer, Lucila Ohno-Machado, Michael E. Matheny |
AMIA | 5 |
| 2015 | Improving Radiology Procedure Identification for Inferior Vena Cava (IVC) Filters using EHR Text
Dalia A. Mobarek, Najeebah Bade, Benjamin Viernes, Patricia Nechodom, Frederick R. Rickles, Scott L. DuVall |
AMIA | 6 |
| 2015 | Building custom lexicon for a large number of related concepts using templates
Olga V. Patterson, Matthew Freiberg, Scott L. DuVall |
AMIA | 3 |
| 2014 | Rapid NLP Development with Leo
Ryan Cornia, Olga V. Patterson, Thomas Ginter, Scott L. DuVall |
AMIA | 4 |
| 2014 | Sophia: An Expedient UMLS Concept Extraction Annotator
Guy Divita, Qing T. Zeng, Adi V. Gundlapalli, Scott L. DuVall, Jonathan R. Nebeker, Matthew H. Samore |
AMIA | 4 |
| 2014 | Check it with Chex: A Validation Tool for Iterative NLP Development
Scott L. DuVall, Ryan Cornia, Tyler Forbush, Corinne Halls, Olga V. Patterson |
AMIA | 1 |
| 2014 | Machine Learning Made Easy with Sherlock
Thomas Ginter, Olga V. Patterson, Ryan Cornia, Scott L. DuVall |
AMIA | 4 |
| 2014 | Automatic Engine for Mapping Mycobacteriology Reports to SNOMED-CT
Olga V. Patterson, Scott D. Nelson, Makoto Jones, Kimberly Findley, Kevin L. Winthrop, Kevin P. Fennelly, Scott L. DuVall |
AMIA | 7 |
| 2013 | Using Natural Language Processing on the Free Text of Clinical Documents to Screen for Evidence of Homelessness Among US Veterans
Adi V. Gundlapalli, Marjorie Carter, Miland N. Palmer, Thomas Ginter, Andrew Redd, Steve Pickard, Shuying Shen, Brett R. South, Guy Divita, Scott L. DuVall, Thien M. Nguyen, Leonard W. D'Avolio, Matthew H. Samore |
AMIA | 10 |
| 2013 | Improving performance of natural language processing part-of-speech tagging on clinical narratives through domain adaptationabstractOBJECTIVE: Natural language processing (NLP) tasks are commonly decomposed into subtasks, chained together to form processing pipelines. The residual error produced in these subtasks propagates, adversely affecting the end objectives. Limited availability of annotated clinical data remains a barrier to reaching state-of-the-art operating characteristics using statistically based NLP tools in the clinical domain. Here we explore the unique linguistic constructions of clinical texts and demonstrate the loss in operating characteristics when out-of-the-box part-of-speech (POS) tagging tools are applied to the clinical domain. We test a domain adaptation approach integrating a novel lexical-generation probability rule used in a transformation-based learner to boost POS performance on clinical narratives. METHODS: Two target corpora from independent healthcare institutions were constructed from high frequency clinical narratives. Four leading POS taggers with their out-of-the-box models trained from general English and biomedical abstracts were evaluated against these clinical corpora. A high performing domain adaptation method, Easy Adapt, was compared to our newly proposed method ClinAdapt. RESULTS: The evaluated POS taggers drop in accuracy by 8.5-15% when tested on clinical narratives. The highest performing tagger reports an accuracy of 88.6%. Domain adaptation with Easy Adapt reports accuracies of 88.3-91.0% on clinical texts. ClinAdapt reports 93.2-93.9%. CONCLUSIONS: ClinAdapt successfully boosts POS tagging performance through domain adaptation requiring a modest amount of annotated clinical data. Improving the performance of critical NLP subtasks is expected to reduce pipeline error propagation leading to better overall results on complex processing tasks. Jeffrey P. Ferraro, Hal Daumé III, Scott L. DuVall, Wendy W. Chapman, Henk Harkema, Peter J. Haug |
J. Am. Medical Informatics Assoc. | 3 |
| 2012 | JMX Analysis Module: Multi-thread aggregate NLP performance monitoring
Ryan Cornia, Olga V. Patterson, Scott L. DuVall |
AMIA | 3 |
| 2012 | CASPR: Friendly Annotation Management
Tyler Forbush, Brad Adams, Shuying Shen, Brett R. South, Jonathan R. Nebeker, Scott L. DuVall |
AMIA | 6 |
| 2012 | Project Management and Coordination: Selecting Communication Tools for Multi-Site, Multidisciplinary Collaboration
James Potter, Corinne Halls, Beniel Malohi, Scott L. DuVall |
AMIA | 4 |
| 2012 | Evaluation and Visualization of Human Annotator Learning Patterns
Ying Suo, Shuying Shen, Scott L. DuVall, Özlem Uzuner, Brett R. South |
AMIA | 3 |
| 2012 | Evaluation of record linkage between a large healthcare provider and the Utah Population DatabaseabstractOBJECTIVE: Electronically linked datasets have become an important part of clinical research. Information from multiple sources can be used to identify comorbid conditions and patient outcomes, measure use of healthcare services, and enrich demographic and clinical variables of interest. Innovative approaches for creating research infrastructure beyond a traditional data system are necessary. MATERIALS AND METHODS: Records from a large healthcare system's enterprise data warehouse (EDW) were linked to a statewide population database, and a master subject index was created. The authors evaluate the linkage, along with the impact of missing information in EDW records and the coverage of the population database. The makeup of the EDW and population database provides a subset of cancer records that exist in both resources, which allows a cancer-specific evaluation of the linkage. RESULTS: About 3.4 million records (60.8%) in the EDW were linked to the population database with a minimum accuracy of 96.3%. It was estimated that approximately 24.8% of target records were absent from the population database, which enabled the effect of the amount and type of information missing from a record on the linkage to be estimated. However, 99% of the records from the oncology data mart linked; they had fewer missing fields and this correlated positively with the number of patient visits. DISCUSSION AND CONCLUSION: A general-purpose research infrastructure was created which allows disease-specific cohorts to be identified. The usefulness of creating an index between institutions is that it allows each institution to maintain control and confidentiality of their own information. Scott L. DuVall, Alison M. Fraser, Kerry Rowe, Alun Thomas, Geraldine P. Mineau |
J. Am. Medical Informatics Assoc. | 1 |
| 2012 | Automated extraction of ejection fraction for quality measurement using regular expressions in Unstructured Information Management Architecture (UIMA) for heart failureabstractOBJECTIVES: Left ventricular ejection fraction (EF) is a key component of heart failure quality measures used within the Department of Veteran Affairs (VA). Our goals were to build a natural language processing system to extract the EF from free-text echocardiogram reports to automate measurement reporting and to validate the accuracy of the system using a comparison reference standard developed through human review. This project was a Translational Use Case Project within the VA Consortium for Healthcare Informatics. MATERIALS AND METHODS: We created a set of regular expressions and rules to capture the EF using a random sample of 765 echocardiograms from seven VA medical centers. The documents were randomly assigned to two sets: a set of 275 used for training and a second set of 490 used for testing and validation. To establish the reference standard, two independent reviewers annotated all documents in both sets; a third reviewer adjudicated disagreements. RESULTS: System test results for document-level classification of EF of <40% had a sensitivity (recall) of 98.41%, a specificity of 100%, a positive predictive value (precision) of 100%, and an F measure of 99.2%. System test results at the concept level had a sensitivity of 88.9% (95% CI 87.7% to 90.0%), a positive predictive value of 95% (95% CI 94.2% to 95.9%), and an F measure of 91.9% (95% CI 91.2% to 92.7%). DISCUSSION: An EF value of <40% can be accurately identified in VA echocardiogram reports. CONCLUSIONS: An automated information extraction system can be used to accurately extract EF for quality measurement. Jennifer H. Garvin, Scott L. DuVall, Brett R. South, Bruce E. Bray, Daniel Bolton, Julia Heavirland, Steve Pickard, Paul Heidenreich, Shuying Shen, Charlene R. Weir, Matthew H. Samore, Mary K. Goldstein |
J. Am. Medical Informatics Assoc. | 2 |
| 2012 | Leveraging Social Bookmarks from Partially Tagged Corpus for Improved Web Page ClusteringabstractAutomatic clustering of Web pages helps a number of information retrieval tasks, such as improving user interfaces, collection clustering, introducing diversity in search results, etc. Typically, Web page clustering algorithms use only features extracted from the page-text. However, the advent of social-bookmarking Web sites, such as StumbleUpon.com and Delicious.com, has led to a huge amount of user-generated content such as the social tag information that is associated with the Web pages. In this article, we present a subspace based feature extraction approach that leverages the social tag information to complement the page-contents of a Web page for extracting beter features, with the goal of improved clustering performance. In our approach, we consider page-text and tags as two separate views of the data, and learn a shared subspace that maximizes the correlation between the two views. Any clustering algorithm can then be applied in this subspace. We then present an extension that allows our approach to be applicable even if the Web page corpus is only partially tagged, that is, when the social tags are present for not all, but only for a small number of Web pages. We compare our subspace based approach with a number of baselines that use tag information in various other ways, and show that the subspace based approach leads to improved performance on the Web page clustering task. We also discuss some possible future work including an active learning extension that can help in choosing which Web pages to get tags for, if we only can get the social tags for only a small number of Web pages. Anusua Trivedi, Piyush Rai, Hal Daumé III, Scott L. DuVall |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2011 | Active Supervised Domain Adaptation
Avishek Saha, Piyush Rai, Hal Daumé III, Suresh Venkatasubramanian, Scott L. DuVall |
ECML/PKDD (3) | 5 |
| 2011 | 2010 i2b2/VA challenge on concepts, assertions, and relations in clinical textabstractThe 2010 i2b2/VA Workshop on Natural Language Processing Challenges for Clinical Records presented three tasks: a concept extraction task focused on the extraction of medical concepts from patient reports; an assertion classification task focused on assigning assertion types for medical problem concepts; and a relation classification task focused on assigning relation types that hold between medical problems, tests, and treatments. i2b2 and the VA provided an annotated reference standard corpus for the three tasks. Using this reference standard, 22 systems were developed for concept extraction, 21 for assertion classification, and 16 for relation classification. These systems showed that machine learning approaches could be augmented with rule-based systems to determine concepts, assertions, and relations. Depending on the task, the rule-based systems can either provide input for machine learning or post-process the output of machine learning. Ensembles of classifiers, information from unlabeled data, and external knowledge sources can help when the training data are inadequate. Özlem Uzuner, Brett R. South, Shuying Shen, Scott L. DuVall |
J. Am. Medical Informatics Assoc. | 4 |
| 2010 | Extending the Fellegi-Sunter probabilistic record linkage method for approximate field comparators
Scott L. DuVall, Richard A. Kerber, Alun Thomas |
J. Biomed. Informatics | 1 |
| 2006 | Academic Podcasting: Quality Media Delivery
Jacob S. Tripp, Scott L. DuVall, Derek L. Cowan, Aaron W. C. Kamauu |
AMIA | 2 |