Sarah E. Schwarm

dblp:53/3743 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
0since 2021 · last 2005
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 74% Information extraction and text analysis · 20% Speech recognition and synthesis · 5%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › text classification
readability assessment
0.112005
Reading Level Assessment Using Support Vector Machines and Statistical Language Models · ACL 2005
Natural language and speech › Language models and text generation › language modeling
statistical language modeling
0.112005
Reading Level Assessment Using Support Vector Machines and Statistical Language Models · ACL 2005
Natural language and speech › Language models and text generation › neural language model
adaptive language models
0.012004
Adaptive language modeling with varied sources to cover new vocabulary items · IEEE Trans. Speech Audio Process. 2004
Natural language and speech › Language models and text generation
language modeling
0.012004
Adaptive language modeling with varied sources to cover new vocabulary items · IEEE Trans. Speech Audio Process. 2004
Natural language and speech › Language models and text generation › language modeling
n-gram language model
0.012004
Adaptive language modeling with varied sources to cover new vocabulary items · IEEE Trans. Speech Audio Process. 2004
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › multi-speaker speech recognition
meeting transcription
0.012004
Adaptive language modeling with varied sources to cover new vocabulary items · IEEE Trans. Speech Audio Process. 2004

Methods — techniques the papers use, named apart from their topics

support vector machine · 0.1statistical language model · 0.1mixture n-gram models · 0.0
YearPublicationVenuePosition
2005 Reading Level Assessment Using Support Vector Machines and Statistical Language Models
abstract
Reading proficiency is a fundamental component of language competency. However, finding topical texts at an appropriate reading level for foreign and second language learners is a challenge for teachers. This task can be addressed with natural language processing technology to assess reading level. Existing measures of reading level are not well suited to this task, but previous work and our own pilot experiments have shown the benefit of using statistical language models. In this paper, we also use support vector machines to combine features from traditional reading level measures, statistical language models, and other language processing tools to produce a better method of assessing reading level.
Sarah E. Schwarm, Mari Ostendorf
ACL1
2004 Detecting Structural Metadata with Decision Trees and Transformation-Based Learning
Joungbum Kim, Sarah E. Schwarm, Mari Ostendorf
HLT-NAACL2
2004 Adaptive language modeling with varied sources to cover new vocabulary items
abstract
N-gram language modeling typically requires large quantities of in-domain training data, i.e., data that matches the task in both topic and style. For conversational speech applications, particularly meeting transcription, obtaining large volumes of speech transcripts is often unrealistic; topics change frequently and collecting conversational-style training data is time-consuming and expensive. In particular, new topics introduce new vocabulary items which are not included in existing models. In this work, we use a variety of data sources (reflecting different sizes and styles), combined using mixture n-gram models. We study the impact of the different sources on vocabulary expansion and recognition accuracy, and investigate possible indicators of the usefulness of a data source. For the task of recognizing meeting speech, we obtain a 9% relative reduction in the overall word error rate and a 61% relative reduction in the word error rate for "new" words added to the vocabulary over a baseline language model trained from general conversational speech data.
Sarah E. Schwarm, Ivan Bulyko, Mari Ostendorf
IEEE Trans. Speech Audio Process.1
2003 Making connections: using classroom assessment to elicit students' prior knowledge and construction of concepts
abstract
Students bring prior knowledge and experiences to the classroom. According to the constructivist learning theory, students incorporate new knowledge into their existing knowledge frameworks. We used Classroom Assessment Techniques in an information technology course to elicit the construction of knowledge process. We found that CATs and instructor feedback can help shape and reveal this construction process. For example, responses to CATs revealed students' understandings of variables, digital representation, and iteration in the information technology domain. Some students claimed that the CATs helped put new ideas into their own words and helped them simplify the concepts.
Sarah E. Schwarm, Tammy VanDeGrift
ITiCSE1
2002 Text normalization with varied data sources for conversational speech language modeling
abstract
Collecting sufficient language model training data for good speech recognition performance in a new domain is often difficult. However, there may be other sources of data that are matched in terms of topic or style, if not both. This paper looks at the use of text normalization tools to make these data more suitable for language model training, in conjunction with mixture models to combine data from different sources. We specifically address the task of recognizing meeting speech, showing a small reduction in word error rate over a baseline language model trained from conversational speech data.
Sarah E. Schwarm, Mari Ostendorf
ICASSP1