VLDB 2026 Research / reviewers in the wild / expert
Sarah E. Schwarm
dblp:53/3743
· DBLP profile ↗
5ranked-venue papers
4as first author
0since 2021 · last 2005
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 74% Information extraction and text analysis · 20% Speech recognition and synthesis · 5% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › text classification
readability assessment |
0.1 | 1 | 2005 | Reading Level Assessment Using Support Vector Machines and Statistical Language Models · ACL 2005 |
Natural language and speech › Language models and text generation › language modeling
statistical language modeling |
0.1 | 1 | 2005 | Reading Level Assessment Using Support Vector Machines and Statistical Language Models · ACL 2005 |
Natural language and speech › Language models and text generation › neural language model
adaptive language models |
0.0 | 1 | 2004 | Adaptive language modeling with varied sources to cover new vocabulary items · IEEE Trans. Speech Audio Process. 2004 |
Natural language and speech › Language models and text generation
language modeling |
0.0 | 1 | 2004 | Adaptive language modeling with varied sources to cover new vocabulary items · IEEE Trans. Speech Audio Process. 2004 |
Natural language and speech › Language models and text generation › language modeling
n-gram language model |
0.0 | 1 | 2004 | Adaptive language modeling with varied sources to cover new vocabulary items · IEEE Trans. Speech Audio Process. 2004 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › multi-speaker speech recognition
meeting transcription |
0.0 | 1 | 2004 | Adaptive language modeling with varied sources to cover new vocabulary items · IEEE Trans. Speech Audio Process. 2004 |
Methods — techniques the papers use, named apart from their topics
support vector machine · 0.1statistical language model · 0.1mixture n-gram models · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2005 | Reading Level Assessment Using Support Vector Machines and Statistical Language ModelsabstractReading proficiency is a fundamental component of language competency. However, finding topical texts at an appropriate reading level for foreign and second language learners is a challenge for teachers. This task can be addressed with natural language processing technology to assess reading level. Existing measures of reading level are not well suited to this task, but previous work and our own pilot experiments have shown the benefit of using statistical language models. In this paper, we also use support vector machines to combine features from traditional reading level measures, statistical language models, and other language processing tools to produce a better method of assessing reading level. Sarah E. Schwarm, Mari Ostendorf |
ACL | 1 |
| 2004 | Detecting Structural Metadata with Decision Trees and Transformation-Based Learning
Joungbum Kim, Sarah E. Schwarm, Mari Ostendorf |
HLT-NAACL | 2 |
| 2004 | Adaptive language modeling with varied sources to cover new vocabulary itemsabstractN-gram language modeling typically requires large quantities of in-domain training data, i.e., data that matches the task in both topic and style. For conversational speech applications, particularly meeting transcription, obtaining large volumes of speech transcripts is often unrealistic; topics change frequently and collecting conversational-style training data is time-consuming and expensive. In particular, new topics introduce new vocabulary items which are not included in existing models. In this work, we use a variety of data sources (reflecting different sizes and styles), combined using mixture n-gram models. We study the impact of the different sources on vocabulary expansion and recognition accuracy, and investigate possible indicators of the usefulness of a data source. For the task of recognizing meeting speech, we obtain a 9% relative reduction in the overall word error rate and a 61% relative reduction in the word error rate for "new" words added to the vocabulary over a baseline language model trained from general conversational speech data. Sarah E. Schwarm, Ivan Bulyko, Mari Ostendorf |
IEEE Trans. Speech Audio Process. | 1 |
| 2003 | Making connections: using classroom assessment to elicit students' prior knowledge and construction of conceptsabstractStudents bring prior knowledge and experiences to the classroom. According to the constructivist learning theory, students incorporate new knowledge into their existing knowledge frameworks. We used Classroom Assessment Techniques in an information technology course to elicit the construction of knowledge process. We found that CATs and instructor feedback can help shape and reveal this construction process. For example, responses to CATs revealed students' understandings of variables, digital representation, and iteration in the information technology domain. Some students claimed that the CATs helped put new ideas into their own words and helped them simplify the concepts. Sarah E. Schwarm, Tammy VanDeGrift |
ITiCSE | 1 |
| 2002 | Text normalization with varied data sources for conversational speech language modelingabstractCollecting sufficient language model training data for good speech recognition performance in a new domain is often difficult. However, there may be other sources of data that are matched in terms of topic or style, if not both. This paper looks at the use of text normalization tools to make these data more suitable for language model training, in conjunction with mixture models to combine data from different sources. We specifically address the task of recognizing meeting speech, showing a small reduction in word error rate over a baseline language model trained from conversational speech data. Sarah E. Schwarm, Mari Ostendorf |
ICASSP | 1 |