EDBT 2026 Demo / reviewers in the wild / expert
Stephen T. Wu
dblp:117/6622 · also Stephen Tze-Inn Wu
· DBLP profile ↗
25ranked-venue papers
15as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 9 first-authorArtificial intelligence and machine learning · 6 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Machine translation · 61% Language models and text generation · 39% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 67% Empirical software engineering · 33% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation › statistical machine translation
phrase-based translation |
0.1 | 1 | 2011 | Incremental Syntactic Language Models for Phrase-based Translation · ACL 2011 |
Natural language and speech › Machine translation
statistical machine translation |
0.1 | 1 | 2011 | Incremental Syntactic Language Models for Phrase-based Translation · ACL 2011 |
Natural language and speech › Language models and text generation › language modeling
syntactic language model |
0.1 | 1 | 2011 | Incremental Syntactic Language Models for Phrase-based Translation · ACL 2011 |
Compilers and program optimization › parsing
incremental parsing |
0.1 | 1 | 2010 | Complexity Metrics in an Incremental Right-Corner Parser · ACL 2010 |
Compilers and program optimization
parsing |
0.1 | 1 | 2010 | Complexity Metrics in an Incremental Right-Corner Parser · ACL 2010 |
Empirical software engineering › software metrics
software complexity metrics |
0.1 | 1 | 2010 | Complexity Metrics in an Incremental Right-Corner Parser · ACL 2010 |
Natural language and speech › Language models and text generation
language modeling |
0.0 | 1 | 2011 | Incremental Syntactic Language Models for Phrase-based Translation · ACL 2011 |
Methods — techniques the papers use, named apart from their topics
syntactic language model · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Modeling asynchronous event sequences with RNNs
Stephen T. Wu, Sijia Liu 0002, Sunghwan Sohn, Sungrim Moon, Chung-Il Wi, Young J. Juhn |
J. Biomed. Informatics | 1 |
| 2017 | Intrainstitutional EHR collections for patient-level information retrievalabstractResearch in clinical information retrieval has long been stymied by the lack of open resources. However, both clinical information retrieval research innovation and legitimate privacy concerns can be served by the creation of intrainstitutional, fully protected resources. In this article, we provide some principles and tools for information retrieval resource‐building in the unique problem setting of patient‐level information retrieval, following the tradition of the Cranfield paradigm. We further include an analysis of parallel information retrieval resources at Oregon Health & Science University and Mayo Clinic that were built on these principles. Stephen T. Wu, Sijia Liu 0002, Yanshan Wang, Tamara Timmons, Harsha Uppili, Steven Bedrick, William R. Hersh |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2016 | Development of Test Topics for Cohort Identification
Tamara Timmons, Stephen T. Wu, William R. Hersh |
AMIA | 2 |
| 2016 | A Part-Of-Speech Weighting Scheme for Clinical Information Retrieval
Yanshan Wang, Stephen T. Wu, Dingcheng Li |
AMIA | 2 |
| 2016 | Restoring line breaks in Epic-derived clinical notes
Stephen T. Wu, Allison Sliter, Meikun Wang, Tamara Timmons, Steven Bedrick |
AMIA | 1 |
| 2016 | Probabilistic Population-level Modeling of Disease Event Timelines
Stephen T. Wu, Yanshan Wang, Sunghwan Sohn, Chung-Il Wi, Elizabeth A. Krusemark, Young J. Juhn |
AMIA | 1 |
| 2016 | On Developing Resources for Patient-level Information Retrieval
Stephen T. Wu, Tamara Timmons, Amy Yates, Meikun Wang, Steven Bedrick, William R. Hersh |
LREC | 1 |
| 2016 | Staggered NLP-assisted refinement for Clinical Annotations of Chronic Disease Events
Stephen T. Wu, Chung-Il Wi, Sunghwan Sohn, Young J. Juhn |
LREC | 1 |
| 2016 | A Part-Of-Speech term weighting scheme for biomedical information retrieval
Yanshan Wang, Stephen T. Wu, Dingcheng Li, Saeed Mehrabi 0003 |
J. Biomed. Informatics | 2 |
| 2014 | Layered Spaces for Clinical Information Retrieval
Stephen T. Wu, Dingcheng Li, James J. Masanz |
AMIA | 1 |
| 2014 | Research and applications: Patient-level temporal aggregation for text-based asthma status ascertainmentabstractOBJECTIVE: To specify the problem of patient-level temporal aggregation from clinical text and introduce several probabilistic methods for addressing that problem. The patient-level perspective differs from the prevailing natural language processing (NLP) practice of evaluating at the term, event, sentence, document, or visit level. METHODS: We utilized an existing pediatric asthma cohort with manual annotations. After generating a basic feature set via standard clinical NLP methods, we introduce six methods of aggregating time-distributed features from the document level to the patient level. These aggregation methods are used to classify patients according to their asthma status in two hypothetical settings: retrospective epidemiology and clinical decision support. RESULTS: In both settings, solid patient classification performance was obtained with machine learning algorithms on a number of evidence aggregation methods, with Sum aggregation obtaining the highest F1 score of 85.71% on the retrospective epidemiological setting, and a probability density function-based method obtaining the highest F1 score of 74.63% on the clinical decision support setting. Multiple techniques also estimated the diagnosis date (index date) of asthma with promising accuracy. DISCUSSION: The clinical decision support setting is a more difficult problem. We rule out some aggregation methods rather than determining the best overall aggregation method, since our preliminary data set represented a practical setting in which manually annotated data were limited. CONCLUSION: Results contrasted the strengths of several aggregation algorithms in different settings. Multiple approaches exhibited good patient classification performance, and also predicted the timing of estimates with reasonable accuracy. Stephen T. Wu, Young J. Juhn, Sunghwan Sohn |
J. Am. Medical Informatics Assoc. | 1 |
| 2014 | Using large clinical corpora for query expansion in text-based cohort identification
Dongqing Zhu, Stephen T. Wu, Ben Carterette |
J. Biomed. Informatics | 2 |
| 2013 | Negation's Not Solved: Reconsidering Negation Annotation and Evaluation
Stephen T. Wu, Timothy A. Miller, James J. Masanz, Matthew Coarr, David Carrell, Scott R. Halgrim, David Harris 0004, Cheryl Clark |
AMIA | 1 |
| 2012 | Towards a semantic lexicon for clinical natural language processing
Stephen T. Wu, Dingcheng Li, Siddhartha Jonnalagadda, Sunghwan Sohn, Kavishwar B. Wagholikar, Peter J. Haug, Stanley M. Huff, Christopher G. Chute |
AMIA | 2 |
| 2012 | Asthma Status Identification with Natural Language Processing
Stephen T. Wu, Young J. Juhn, Sunghwan Sohn, K. E. Ravikumar, Kavishwar B. Wagholikar, Siddhartha Jonnalagadda |
AMIA | 1 |
| 2012 | Tracking Immigrant Health with Natural Language Processing
Stephen T. Wu, Mark Wieland, Vinod Kaggal, Christopher G. Chute |
AMIA | 1 |
| 2012 | Coreference analysis in clinical notes: a multi-pass sieve with alternate anaphora resolution modulesabstractOBJECTIVE: This paper describes the coreference resolution system submitted by Mayo Clinic for the 2011 i2b2/VA/Cincinnati shared task Track 1C. The goal of the task was to construct a system that links the markables corresponding to the same entity. MATERIALS AND METHODS: The task organizers provided progress notes and discharge summaries that were annotated with the markables of treatment, problem, test, person, and pronoun. We used a multi-pass sieve algorithm that applies deterministic rules in the order of preciseness and simultaneously gathers information about the entities in the documents. Our system, MedCoref, also uses a state-of-the-art machine learning framework as an alternative to the final, rule-based pronoun resolution sieve. RESULTS: The best system that uses a multi-pass sieve has an overall score of 0.836 (average of B(3), MUC, Blanc, and CEAF F score) for the training set and 0.843 for the test set. DISCUSSION: A supervised machine learning system that typically uses a single function to find coreferents cannot accommodate irregularities encountered in data especially given the insufficient number of examples. On the other hand, a completely deterministic system could lead to a decrease in recall (sensitivity) when the rules are not exhaustive. The sieve-based framework allows one to combine reliable machine learning components with rules designed by experts. CONCLUSION: Using relatively simple rules, part-of-speech information, and semantic type properties, an effective coreference resolution system could be designed. The source code of the system described is available at https://sourceforge.net/projects/ohnlp/files/MedCoref. Siddhartha Jonnalagadda, Dingcheng Li, Sunghwan Sohn, Stephen T. Wu, Kavishwar B. Wagholikar, Manabu Torii |
J. Am. Medical Informatics Assoc. | 4 |
| 2012 | Unified Medical Language System term occurrences in clinical notes: a large-scale corpus analysisabstractOBJECTIVE: To characterise empirical instances of Unified Medical Language System (UMLS) Metathesaurus term strings in a large clinical corpus, and to illustrate what types of term characteristics are generalisable across data sources. DESIGN: Based on the occurrences of UMLS terms in a 51 million document corpus of Mayo Clinic clinical notes, this study computes statistics about the terms' string attributes, source terminologies, semantic types and syntactic categories. Term occurrences in 2010 i2b2/VA text were also mapped; eight example filters were designed from the Mayo-based statistics and applied to i2b2/VA data. RESULTS: For the corpus analysis, negligible numbers of mapped terms in the Mayo corpus had over six words or 55 characters. Of source terminologies in the UMLS, the Consumer Health Vocabulary and Systematized Nomenclature of Medicine-Clinical Terms (SNOMED-CT) had the best coverage in Mayo clinical notes at 106426 and 94788 unique terms, respectively. Of 15 semantic groups in the UMLS, seven groups accounted for 92.08% of term occurrences in Mayo data. Syntactically, over 90% of matched terms were in noun phrases. For the cross-institutional analysis, using five example filters on i2b2/VA data reduces the actual lexicon to 19.13% of the size of the UMLS and only sees a 2% reduction in matched terms. CONCLUSION: The corpus statistics presented here are instructive for building lexicons from the UMLS. Features intrinsic to Metathesaurus terms (well formedness, length and language) generalise easily across clinical institutions, but term frequencies should be adapted with caution. The semantic groups of mapped terms may differ slightly from institution to institution, but they differ greatly when moving to the biomedical literature domain. Stephen T. Wu, Dingcheng Li, Cui Tao, Mark A. Musen, Christopher G. Chute, Nigam H. Shah |
J. Am. Medical Informatics Assoc. | 1 |
| 2012 | Enhancing clinical concept extraction with distributional semantics
Siddhartha Jonnalagadda, Trevor Cohen, Stephen T. Wu, Graciela Gonzalez-Hernandez |
J. Biomed. Informatics | 3 |
| 2011 | Incremental Syntactic Language Models for Phrase-based Translation
Lane Schwartz, Chris Callison-Burch, William Schuler, Stephen T. Wu |
ACL | 4 |
| 2010 | Complexity Metrics in an Incremental Right-Corner Parser
Stephen T. Wu, Asaf Bachrach, Carlos Cardenas, William Schuler |
ACL | 1 |
| 2009 | A Framework for Fast Incremental Interpretation during Speech DecodingabstractThis article describes a framework for incorporating referential semantic information from a world model or ontology directly into a probabilistic language model of the sort commonly used in speech recognition, where it can be probabilistically weighted together with phonological and syntactic factors as an integral part of the decoding process. Introducing world model referents into the decoding search greatly increases the search space, but by using a single integrated phonological, syntactic, and referential semantic language model, the decoder is able to incrementally prune this search based on probabilities associated with these combined contexts. The result is a single unified referential semantic probability model which brings several kinds of context to bear in speech decoding, and performs accurate recognition in real time on large domains in the absence of example in-domain training sentences. William Schuler, Stephen T. Wu, Lane Schwartz |
Comput. Linguistics | 2 |
| 2008 | Referential semantic language modeling for data-poor domainsabstractThis paper describes a referential semantic language model that achieves accurate recognition in user-defined domains with no available domain-specific training corpora. This model is interesting in that, unlike similar recent systems, it exploits context dynamically, using incremental processing and limited stack memory of an HMM-like time series model to constrain search. Stephen T. Wu, Lane Schwartz, William Schuler |
ICASSP | 1 |
| 2008 | Exploiting referential context in spoken language interfaces for data-poor domainsabstractThis paper describes an implementation of a shell-like programming interface that utilizes referential context (that is, information about the current state of an interfaced application) in order to achieve accurate recognition -- even in user-defined domains with no available domain-specific training corpora. The interface incorporates a knowledge of context into its model of syntax, yielding a referential semantic language model. Interestingly, the referential semantic language model exploits context dynamically, unlike other recent systems, by using incremental processing and the limited stack memory of an HMM-like time series model. Stephen T. Wu, Lane Schwartz, William Schuler |
IUI | 1 |
| 2006 | Dynamic evidence models in a DBN phone recognizerabstractThis paper describes an implementation of a discriminative acoustical model – a Conditional Random Field (CRF) – within a Dynamic Bayes Net (DBN) formulation of a Hierarchic Hidden Markov Model (HHMM) phone recognizer. This CRF-DBN topology accounts for phone transition dynamics in conditional probability distributions over random variables associated with observed evidence, and therefore has less need for hidden variable states corresponding to transitions between phones, leaving more hypothesis space available for modeling higher-level linguistic phenomena such syntax and semantics. The model also has the interesting property that it explicitly represents likely formant trajectories and formant targets of modeled phones in its random variable distributions, making it more linguistically transparent than models based on traditional HMMs with conditionally independent evidence variables. Results on the standard TIMIT phone recognition task show this CRF evidence model, even with a relatively simple first-order feature set, is competitive with standard HMMs and DBN variants using static Gaussian mixture models on MFCC features. William Schuler, Timothy A. Miller, Stephen T. Wu, Andrew Exley |
INTERSPEECH | 3 |