EDBT 2026 Demo / reviewers in the wild / expert
William Schuler
dblp:21/41
· DBLP profile ↗
47ranked-venue papers
12as first author
12since 2021 · last 2026
0000-0002-8780-2174ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 44 · 12 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating Humanlike Memory Effects in Transformers Using Item Recognition TasksabstractRecent studies examining cued recall in Transformers have observed that these language models remember information from the beginning or end of a passage more easily than information in the middle, a pattern which is evocative of serial position effects (primacy and recency) observed in human memory.However, while these effects have been documented in humans across a range of memory tasks (e.g., serial recall, free recall, item recognition), it is less clear whether they generalize beyond cued recall in Transformers.We address this limitation of previous work by performing novel behavioral evaluations on Transformers using a simple item recognition paradigm, which we compare against evaluations using cued recall.We find that Transformers show weak or absent recency effects in item recognition, a pattern which differs from human behavior and from Transformers' own behavior in cued recall.A subsequent experiment examines the role of Transformers' architectural biases in producing serial position effects in item recognition and cued recall. Christian Clark 0001, William Schuler |
CoNLL | 2 |
| 2025 | The Impact of Token Granularity on the Predictive Power of Language Model SurprisalabstractWord-by-word language model surprisal is often used to model the incremental processing of human readers, which raises questions about how various choices in language modeling influence its predictive power.One factor that has been overlooked in cognitive modeling is the granularity of subword tokens, which explicitly encodes information about word length and frequency, and ultimately influences the quality of vector representations that are learned.This paper presents experiments that manipulate the token granularity and evaluate its impact on the ability of surprisal to account for processing difficulty of naturalistic text and garden-path constructions.Experiments with naturalistic reading times reveal a substantial influence of token granularity on surprisal, with tokens defined by a vocabulary size of 8,000 resulting in surprisal that is most predictive.In contrast, on gardenpath constructions, language models trained on coarser-grained tokens generally assigned higher surprisal to critical regions, suggesting a greater sensitivity to garden-path effects than previously reported.Taken together, these results suggest a large role of token granularity on the quality of language model surprisal for cognitive modeling. Byung-Doh Oh, William Schuler |
ACL (1) | 2 |
| 2025 | Linear Recency Bias During Training Improves Transformers' Fit to Reading TimesabstractRecent psycholinguistic research has compared human reading times to surprisal estimates from language models to study the factors shaping human sentence processing difficulty. Previous studies have shown a strong fit between surprisal values from Transformers and reading times. However, standard Transformers work with a lossless representation of the entire previous linguistic context, unlike models of human language processing that include memory decay. To bridge this gap, this paper evaluates a modification of the Transformer model that uses ALiBi (Press et al., 2022), a recency bias added to attention scores. Surprisal estimates from a Transformer that includes ALiBi during training and inference show an improved fit to human reading times compared to a standard Transformer baseline. A subsequent analysis of attention heads suggests that ALiBi’s mixture of slopes—which determine the rate of memory decay in each attention head—may play a role in the improvement by helping models with ALiBi to track different kinds of linguistic dependencies. Christian Clark 0001, Byung-Doh Oh, William Schuler |
COLING | 3 |
| 2024 | Categorial Grammar Induction with Stochastic Category SelectionabstractGrammar induction, the task of learning a set of syntactic rules from minimally annotated training data, provides a means of exploring the longstanding question of whether humans rely on innate knowledge to acquire language. Of the various formalisms available for grammar induction, categorial grammars provide an appealing option due to their transparent interface between syntax and semantics. However, to obtain competitive results, previous categorial grammar inducers have relied on shortcuts such as part-of-speech annotations or an ad hoc bias term in the objective function to ensure desirable branching behavior. We present a categorial grammar inducer that eliminates both shortcuts: it learns from raw data, and does not rely on a biased objective function. This improvement is achieved through a novel stochastic process used to select the set of available syntactic categories. On a corpus of English child-directed speech, the model attains a recall-homogeneity of 0.48, a large improvement over previous categorial grammar inducers. Christian Clark 0001, William Schuler |
LREC/COLING | 2 |
| 2024 | Frequency Explains the Inverse Correlation of Large Language Models' Size, Training Data Amount, and Surprisal's Fit to Reading TimesabstractRecent studies have shown that as Transformerbased language models become larger and are trained on very large amounts of data, the fit of their surprisal estimates to naturalistic human reading times degrades.The current work presents a series of analyses showing that word frequency is a key explanatory factor underlying these two trends.First, residual errors from four language model families on four corpora show that the inverse correlation between model size and fit to reading times is the strongest on the subset of least frequent words, which is driven by excessively accurate predictions of larger model variants.Additionally, training dynamics reveal that during later training steps, all model variants learn to predict rare words and that larger model variants do so more accurately, which explains the detrimental effect of both training data amount and model size on fit to reading times.Finally, a feature attribution analysis demonstrates that larger model variants are able to accurately predict rare words based on both an effectively longer context window size as well as stronger local associations compared to smaller model variants.Taken together, these results indicate that Transformer-based language models' surprisal estimates diverge from human-like expectations due to the superhumanly complex associations they learn for predicting rare words. Byung-Doh Oh, Shisen Yue, William Schuler |
EACL (1) | 3 |
| 2024 | Leading Whitespaces of Language Models' Subword Vocabulary Pose a Confound for Calculating Word ProbabilitiesabstractPredictions of word-by-word conditional probabilities from Transformer-based language models are often evaluated to model the incremental processing difficulty of human readers.In this paper, we argue that there is a confound posed by the most common method of aggregating subword probabilities of such language models into word probabilities.This is due to the fact that tokens in the subword vocabulary of most language models have leading whitespaces and therefore do not naturally define stop probabilities of words.We first prove that this can result in distributions over word probabilities that sum to more than one, thereby violating the axiom that P(Ω) = 1.This property results in a misallocation of word-by-word surprisal, where the unacceptability of the end of the current word is incorrectly carried over to the next word.Additionally, this implicit prediction of word boundaries incorrectly models psycholinguistic experiments where human subjects directly observe upcoming word boundaries.We present a simple decoding technique to reaccount the probability of the trailing whitespace into that of the current word, which resolves this confound.Experiments show that this correction reveals lower estimates of garden-path effects in transitive/intransitive sentences and poorer fits to naturalistic reading times. Byung-Doh Oh, William Schuler |
EMNLP | 2 |
| 2023 | Token-wise Decomposition of Autoregressive Language Model Hidden States for Analyzing Model PredictionsabstractWhile there is much recent interest in studying why Transformer-based large language models make predictions the way they do, the complex computations performed within each layer have made their behavior somewhat opaque.To mitigate this opacity, this work presents a linear decomposition of final hidden states from autoregressive language models based on each initial input token, which is exact for virtually all contemporary Transformer architectures.This decomposition allows the definition of probability distributions that ablate the contribution of specific input tokens, which can be used to analyze their influence on model probabilities over a sequence of upcoming words with only one forward pass from the model.Using the change in next-word probability as a measure of importance, this work first examines which context words make the biggest contribution to language model predictions.Regression experiments suggest that Transformer-based language models rely primarily on collocational associations, followed by linguistic factors such as syntactic dependencies and coreference relationships in making next-word predictions.Additionally, analyses using these measures to predict syntactic dependencies and coreferent mention spans show that collocational association and repetitions of the same token largely explain the language models' predictions on these tasks. Byung-Doh Oh, William Schuler |
ACL (1) | 2 |
| 2023 | Bootstrapping a Conversational Guide for Colonoscopy PrepabstractPulkit Arya, Madeleine Bloomquist, Subhankar Chakraborty, Andrew Perrault, William Schuler, Eric Fosler-Lussier, Michael White. Proceedings of the 24th Meeting of the Special Interest Group on Discourse and Dialogue. 2023. Pulkit Arya, Madeleine Bloomquist, Subhankar Chakraborty, Andrew Perrault, William Schuler, Eric Fosler-Lussier, Michael White 0001 |
SIGDIAL | 5 |
| 2023 | Why Does Surprisal From Larger Transformer-Based Language Models Provide a Poorer Fit to Human Reading Times?abstractAbstract This work presents a linguistic analysis into why larger Transformer-based pre-trained language models with more parameters and lower perplexity nonetheless yield surprisal estimates that are less predictive of human reading times. First, regression analyses show a strictly monotonic, positive log-linear relationship between perplexity and fit to reading times for the more recently released five GPT-Neo variants and eight OPT variants on two separate datasets, replicating earlier results limited to just GPT-2 (Oh et al., 2022). Subsequently, analysis of residual errors reveals a systematic deviation of the larger variants, such as underpredicting reading times of named entities and making compensatory overpredictions for reading times of function words such as modals and conjunctions. These results suggest that the propensity of larger Transformer-based models to ‘memorize’ sequences during training makes their surprisal estimates diverge from humanlike expectations, which warrants caution in using pre-trained language models to study human language processing. Byung-Doh Oh, William Schuler |
Trans. Assoc. Comput. Linguistics | 2 |
| 2022 | Entropy- and Distance-Based Predictors From GPT-2 Attention Patterns Predict Reading Times Over and Above GPT-2 SurprisalabstractTransformer-based large language models are trained to make predictions about the next word by aggregating representations of previous tokens through their self-attention mechanism.In Byung-Doh Oh, William Schuler |
EMNLP | 2 |
| 2021 | Surprisal Estimators for Human Reading Times Need Character ModelsabstractByung-Doh Oh, Christian Clark, William Schuler. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Byung-Doh Oh, Christian Clark 0001, William Schuler |
ACL/IJCNLP (1) | 3 |
| 2021 | Depth-Bounded Statistical PCFG Induction as a Model of Human Grammar AcquisitionabstractAbstract This article describes a simple PCFG induction model with a fixed category domain that predicts a large majority of attested constituent boundaries, and predicts labels consistent with nearly half of attested constituent labels on a standard evaluation data set of child-directed speech. The article then explores the idea that the difference between simple grammars exhibited by child learners and fully recursive grammars exhibited by adult learners may be an effect of increasing working memory capacity, where the shallow grammars are constrained images of the recursive grammars. An implementation of these memory bounds as limits on center embedding in a depth-specific transform of a recursive grammar yields a significant improvement over an equivalent but unbounded baseline, suggesting that this arrangement may indeed confer a learning advantage. Lifeng Jin, Lane Schwartz, Finale Doshi-Velez, Timothy A. Miller, William Schuler |
Comput. Linguistics | 5 |
| 2020 | Coreference information guides human expectations during natural readingabstractModels of human sentence processing effort tend to focus on costs associated with retrieving structures and discourse referents from memory (memory-based) and/or on costs associated with anticipating upcoming words and structures based on contextual cues (expectation-based) (Levy, 2008).Although evidence suggests that expectation and memory may play separable roles in language comprehension (Levy et al., 2013), theories of coreference processing have largely focused on memory: how comprehenders identify likely referents of linguistic expressions.In this study, we hypothesize that coreference tracking also informs human expectations about upcoming words, and we test this hypothesis by evaluating the degree to which incremental surprisal measures generated by a novel coreference-aware semantic parser explain human response times in a naturalistic self-paced reading experiment.Results indicate (1) that coreference information indeed guides human expectations and (2) that coreference effects on memory retrieval may exist independently of coreference effects on expectations.Together, these findings suggest that the language processing system exploits coreference information both to retrieve referents from memory and to anticipate upcoming material. Evan Jaffe, Cory Shain, William Schuler |
COLING | 3 |
| 2020 | A Corpus of Encyclopedia Articles with Logical FormsabstractPeople can extract precise, complex logical meanings from text in documents such as tax forms and game rules, but language processing systems lack adequate training and evaluation resources to do these kinds of tasks reliably. This paper describes a corpus of annotated typed lambda calculus translations for approximately 2,000 sentences in Simple English Wikipedia, which is assumed to constitute a broad-coverage domain for precise, complex descriptions. The corpus described in this paper contains a large number of quantifiers and interesting scoping configurations, and is presented specifically as a resource for quantifier scope disambiguation systems, but also more generally as an object of linguistic study. Nathan Rasmussen, William Schuler |
LREC | 2 |
| 2019 | Unsupervised Learning of PCFGs with Normalizing FlowabstractUnsupervised PCFG inducers hypothesize sets of compact context-free rules as explanations for sentences.These models not only provide tools for low-resource languages, but also play an important role in modeling language acquisition (Bannard et al., 2009;Abend et al., 2017).However, current PCFG induction models, using word tokens as input, are unable to incorporate semantics and morphology into induction, and may encounter issues of sparse vocabulary when facing morphologically rich languages.This paper describes a neural PCFG inducer which employs context embeddings (Peters et al., 2018) in a normalizing flow model (Dinh et al., 2015) to extend PCFG induction to use semantic and morphological information 1 .Linguistically motivated similarity penalty and categorical distance constraints are imposed on the inducer as regularization.Experiments show that the PCFG induction model with normalizing flow produces grammars with state-of-the-art accuracy on a variety of different languages.Ablation further shows a positive effect of normalizing flow, context embeddings and proposed regularizers. Lifeng Jin, Finale Doshi-Velez, Timothy A. Miller, Lane Schwartz, William Schuler |
ACL (1) | 5 |
| 2019 | Variance of Average Surprisal: A Better Predictor for Quality of Grammar from Unsupervised PCFG InductionabstractIn unsupervised grammar induction, data likelihood is known to be only weakly correlated with parsing accuracy, especially at convergence after multiple runs.In order to find a better indicator for quality of induced grammars, this paper correlates several linguistically-and psycholinguisticallymotivated predictors to parsing accuracy on a large multilingual grammar induction evaluation data set.Results show that variance of average surprisal (VAS) better correlates with parsing accuracy than data likelihood, and that using VAS instead of data likelihood for model selection provides a significant accuracy boost.Further evidence shows VAS to be a better candidate than data likelihood for predicting word order typology classification.Analyses show that VAS seems to separate content words from function words in natural language grammars, and to better arrange words with different frequencies into separate classes that are more consistent with linguistic theory. Lifeng Jin, William Schuler |
ACL (1) | 2 |
| 2018 | Depth-bounding is effective: Improvements and Evaluation of Unsupervised PCFG InductionabstractThere have been several recent attempts to improve the accuracy of grammar induction systems by bounding the recursive complexity of the induction model (Ponvert et al., 2011;Noji and Johnson, 2016;Shain et al., 2016;Jin et al., 2018).Modern depth-bounded grammar inducers have been shown to be more accurate than early unbounded PCFG inducers, but this technique has never been compared against unbounded induction within the same system, in part because most previous depthbounding models are built around sequence models, the complexity of which grows exponentially with the maximum allowed depth.The present work instead applies depth bounds within a chart-based Bayesian PCFG inducer (Johnson et al., 2007b), where bounding can be switched on and off, and then samples trees with and without bounding.1 Results show that depth-bounding is indeed significantly effective in limiting the search space of the inducer and thereby increasing the accuracy of the resulting parsing model.Moreover, parsing results on English, Chinese and German show that this bounded model with a new inference technique is able to produce parse trees more accurately than or competitively with state-ofthe-art constituency-based grammar induction models. Lifeng Jin, Finale Doshi-Velez, Timothy A. Miller, William Schuler, Lane Schwartz |
EMNLP | 4 |
| 2018 | Deconvolutional time series regression: A technique for modeling temporally diffuse effectsabstractResearchers in computational psycholinguistics frequently use linear models to study time series data generated by human subjects.However, time series may violate the assumptions of these models through temporal diffusion, where stimulus presentation has a lingering influence on the response as the rest of the experiment unfolds.This paper proposes a new statistical model that borrows from digital signal processing by recasting the predictors and response as convolutionally-related signals, using recent advances in machine learning to fit latent impulse response functions (IRFs) of arbitrary shape.A synthetic experiment shows successful recovery of true latent IRFs, and psycholinguistic experiments reveal plausible, replicable, and fine-grained estimates of latent temporal dynamics, with comparable or improved prediction quality to widely-used alternatives. Cory Shain, William Schuler |
EMNLP | 2 |
| 2018 | Test Sets for Chinese Nonlocal Dependency Parsing
Manjuan Duan, William Schuler |
LREC | 2 |
| 2018 | Unsupervised Grammar Induction with Depth-bounded PCFGabstractThere has been recent interest in applying cognitively- or empirically-motivated bounds on recursion depth to limit the search space of grammar induction models (Ponvert et al., 2011; Noji and Johnson, 2016; Shain et al., 2016). This work extends this depth-bounding approach to probabilistic context-free grammar induction (DB-PCFG), which has a smaller parameter space than hierarchical sequence models, and therefore more fully exploits the space reductions of depth-bounding. Results for this model on grammar acquisition from transcribed child-directed speech and newswire text exceed or are competitive with those of other models when evaluated on parse accuracy. Moreover, grammars acquired from this model demonstrate a consistent use of category labels, something which has not been demonstrated by other acquisition models. Lifeng Jin, Finale Doshi-Velez, Timothy A. Miller, William Schuler, Lane Schwartz |
Trans. Assoc. Comput. Linguistics | 4 |
| 2017 | Approximations of Predictive Entropy Correlate with Reading Times
Marten van Schijndel, William Schuler |
CogSci | 2 |
| 2016 | Memory-Bounded Left-Corner Unsupervised Grammar Induction on Child-Directed InputabstractThis paper presents a new memory-bounded left-corner parsing model for unsupervised raw-text syntax induction, using unsupervised hierarchical hidden Markov models (UHHMM). We deploy this algorithm to shed light on the extent to which human language learners can discover hierarchical syntax through distributional statistics alone, by modeling two widely-accepted features of human language acquisition and sentence processing that have not been simultaneously modeled by any existing grammar induction algorithm: (1) a left-corner parsing strategy and (2) limited working memory capacity. To model realistic input to human language learners, we evaluate our system on a corpus of child-directed speech rather than typical newswire corpora. Results beat or closely match those of three competing systems. Cory Shain, William Bryce, Lifeng Jin, Victoria Krakovna, Finale Doshi-Velez, Timothy A. Miller, William Schuler, Lane Schwartz |
COLING | 7 |
| 2015 | A Comparison of Word Similarity Performance Using Explanatory and Non-explanatory TextsabstractVectorial representations of words derived from large current events datasets have been shown to perform well on word similarity tasks.This paper shows vectorial representations derived from substantially smaller explanatory text datasets such as English Wikipedia and Simple English Wikipedia preserve enough lexical semantic information to make these kinds of category judgments with equal or better accuracy. Lifeng Jin, William Schuler |
HLT-NAACL | 2 |
| 2015 | Hierarchic syntax improves reading time predictionabstractPrevious work has debated whether humans make use of hierarchic syntax when processing language (Frank and Bod, 2011;Fossum and Levy, 2012).This paper uses an eye-tracking corpus to demonstrate that hierarchic syntax significantly improves reading time prediction over a strong n-gram baseline.This study shows that an interpolated 5-gram baseline can be made stronger by combining n-gram statistics over entire eye-tracking regions rather than simply using the last n-gram in each region, but basic hierarchic syntactic measures are still able to achieve significant improvements over this improved baseline. Marten van Schijndel, William Schuler |
HLT-NAACL | 2 |
| 2014 | Frequency effects in the processing of unbounded dependencies
Marten van Schijndel, William Schuler, Peter W. Culicover |
CogSci | 2 |
| 2013 | An Analysis of Frequency- and Memory-Based Processing Costs
Marten van Schijndel, William Schuler |
HLT-NAACL | 2 |
| 2012 | Accurate Unbounded Dependency Recovery using Generalized Categorial Grammars
Luan Nguyen, Marten van Schijndel, William Schuler |
COLING | 3 |
| 2011 | A Pronoun Anaphora Resolution System based on Factorial Hidden Markov Models
Dingcheng Li, Timothy A. Miller, William Schuler |
ACL | 3 |
| 2011 | Incremental Syntactic Language Models for Phrase-based Translation
Lane Schwartz, Chris Callison-Burch, William Schuler, Stephen T. Wu |
ACL | 3 |
| 2011 | Effects of Filler-gap Dependencies Working Memory Requirements for Parsing
William Schuler |
CogSci | 1 |
| 2010 | Complexity Metrics in an Incremental Right-Corner Parser
Stephen T. Wu, Asaf Bachrach, Carlos Cardenas, William Schuler |
ACL | 4 |
| 2010 | Broad-Coverage Parsing Using Human-Like Memory ConstraintsabstractHuman syntactic processing shows many signs of taking place within a general-purpose short-term memory. But this kind of memory is known to have a severely constrained storage capacity—possibly constrained to as few as three or four distinct elements. This article describes a model of syntactic processing that operates successfully within these severe constraints, by recognizing constituents in a right-corner transformed representation (a variant of left-corner parsing) and mapping this representation to random variables in a Hierarchic Hidden Markov Model, a factored time-series model which probabilistically models the contents of a bounded memory store over time. Evaluations of the coverage of this model on a large syntactically annotated corpus of English sentences, and the accuracy of a a bounded-memory parsing strategy based on this model, suggest this model may be cognitively plausible. William Schuler, Samir AbdelRahman, Timothy A. Miller, Lane Schwartz |
Comput. Linguistics | 1 |
| 2009 | Positive effects of redundant descriptions in an interactive semantic speech interfaceabstractSpoken language interfaces based on interactive semantic language models allow probabilities for hypothesized words to be conditioned on the semantic interpretation of these words in the context of some interfaced application environment. This conditioning may allow users to avoid recognition errors in an intuitive way, by adding extra, possibly redundant description. This paper evaluates the effect on error reduction of redundant descriptions in an interactive semantic language model. In order to evaluate the effect in natural use, the model is run on rich domains, supporting references to sets of individuals (instead of just individuals themselves) arranged in multiple continuous dimensions (a 2-D floorplan scene). Results of these experiments suggest that an interactive semantic language model allows users to achieve significantly higher recognition accuracy by providing additional redundant spoken description. Lane Schwartz, Luan Nguyen, Andrew Exley, William Schuler |
IUI | 4 |
| 2009 | Positive Results for Parsing with a Bounded Stack using a Model-Based Right-Corner Transform
William Schuler |
HLT-NAACL | 1 |
| 2009 | A Framework for Fast Incremental Interpretation during Speech DecodingabstractThis article describes a framework for incorporating referential semantic information from a world model or ontology directly into a probabilistic language model of the sort commonly used in speech recognition, where it can be probabilistically weighted together with phonological and syntactic factors as an integral part of the decoding process. Introducing world model referents into the decoding search greatly increases the search space, but by using a single integrated phonological, syntactic, and referential semantic language model, the decoder is able to incrementally prune this search based on probabilities associated with these combined contexts. The result is a single unified referential semantic probability model which brings several kinds of context to bear in speech decoding, and performs accurate recognition in real time on large domains in the absence of example in-domain training sentences. William Schuler, Stephen T. Wu, Lane Schwartz |
Comput. Linguistics | 1 |
| 2008 | A Syntactic Time-Series Model for Parsing Fluent and Disfluent Speech
Timothy A. Miller, William Schuler |
COLING | 2 |
| 2008 | Toward a Psycholinguistically-Motivated Model of Language Processing
William Schuler, Samir AbdelRahman, Timothy A. Miller, Lane Schwartz |
COLING | 1 |
| 2008 | Referential semantic language modeling for data-poor domainsabstractThis paper describes a referential semantic language model that achieves accurate recognition in user-defined domains with no available domain-specific training corpora. This model is interesting in that, unlike similar recent systems, it exploits context dynamically, using incremental processing and limited stack memory of an HMM-like time series model to constrain search. Stephen T. Wu, Lane Schwartz, William Schuler |
ICASSP | 3 |
| 2008 | Exploiting referential context in spoken language interfaces for data-poor domainsabstractThis paper describes an implementation of a shell-like programming interface that utilizes referential context (that is, information about the current state of an interfaced application) in order to achieve accurate recognition -- even in user-defined domains with no available domain-specific training corpora. The interface incorporates a knowledge of context into its model of syntax, yielding a referential semantic language model. Interestingly, the referential semantic language model exploits context dynamically, unlike other recent systems, by using incremental processing and the limited stack memory of an HMM-like time series model. Stephen T. Wu, Lane Schwartz, William Schuler |
IUI | 3 |
| 2007 | Elements of a spoken language programming interface for robotsabstractIn many settings, such as home care or mobile environments, demands on users' attention, or users' anticipated level of formal training, or other on-site conditions will make standard keyboard-and monitor-based robot programming interfaces impractical. In such cases, a spoken language interface may be preferable. However, the open-ended task of programming a machine is very different from the sort of closed-vocabulary, data-rich applications (e.g. call routing) for which most speaker-independent spoken language interfaces are designed. This paper will describe some of the challenges of designing a spoken language programming interface for robots, and will present an approach that uses these semantic-level resources as extensively as possible in order to address these challenges. Timothy A. Miller, Andrew Exley, William Schuler |
HRI | 3 |
| 2006 | Dynamic evidence models in a DBN phone recognizerabstractThis paper describes an implementation of a discriminative acoustical model – a Conditional Random Field (CRF) – within a Dynamic Bayes Net (DBN) formulation of a Hierarchic Hidden Markov Model (HHMM) phone recognizer. This CRF-DBN topology accounts for phone transition dynamics in conditional probability distributions over random variables associated with observed evidence, and therefore has less need for hidden variable states corresponding to transitions between phones, leaving more hypothesis space available for modeling higher-level linguistic phenomena such syntax and semantics. The model also has the interesting property that it explicitly represents likely formant trajectories and formant targets of modeled phones in its random variable distributions, making it more linguistically transparent than models based on traditional HMMs with conditionally independent evidence variables. Results on the standard TIMIT phone recognition task show this CRF evidence model, even with a relatively simple first-order feature set, is competitive with standard HMMs and DBN variants using static Gaussian mixture models on MFCC features. William Schuler, Timothy A. Miller, Stephen T. Wu, Andrew Exley |
INTERSPEECH | 1 |
| 2005 | Integrating denotational meaning into a DBN language modelabstractThis paper describes a dynamic Bayes net (DBN) language model which allows recognition decisions to be conditioned on features of entities in some environment, to which hypothesized directives might refer. The accuracy of this model is then evaluated on spoken directives in various domains. 1. William Schuler, Timothy A. Miller |
INTERSPEECH | 1 |
| 2003 | Using Model-Theoretic Semantic Interpretation to Guide Statistical Parsing and Word Recognition in a Spoken Language InterfaceabstractThis paper describes an extension of the semantic grammars used in conventional statistical spoken language interfaces to allow the probabilities of derived analyses to be conditioned on the meanings or denotations of input utterances in the context of an interface's underlying application environment or world model. Since these denotations will be used to guide disambiguation in interactive applications, they must be efficiently shared among the many possible analyses that may be assigned to an input utterance. This paper therefore presents a formal restriction on the scope of variables in a semantic grammar which guarantees that the denotations of all possible analyses of an input utterance can be calculated in polynomial time, without undue constraints on the expressivity of the derived semantics. Empirical tests show that this model-theoretic interpretation yields a statistically significant improvement on standard measures of parsing accuracy over a baseline grammar not conditioned on denotations. William Schuler |
ACL | 1 |
| 2002 | Interleaved Semantic Interpretation in Environment-based Parsing
William Schuler |
COLING | 1 |
| 2001 | Computational Properties of Environment-based DisambiguationabstractThe standard pipeline approach to semantic processing, in which sentences are morphologically and syntactically resolved to a single tree before they are interpreted, is a poor fit for applications such as natural language interfaces. This is because the environment information, in the form of the objects and events in the application's runtime environment, cannot be used to inform parsing decisions unless the input sentence is semantically analyzed, but this does not occur until after parsing in the single-tree semantic architecture. This paper describes the computational properties of an alternative architecture, in which semantic analysis is performed on all possible interpretations during parsing, in polynomial time. William Schuler |
ACL | 1 |
| 2000 | Multi-Component TAG and Notions of Formal PowerabstractThis paper presents a restricted version of Set-Local Multi-Component TAGs (Weir, 1988) which retains the strong generative capacity of Tree-Local Multi-Component TAG (i.e. produces the same derived structures) but has a greater derivational generative capacity (i.e. can derive those structures in more ways). This formalism is then applied as a framework for integrating dependency and constituency based linguistic representations. William Schuler, David Chiang 0001, Mark Dras |
ACL | 1 |
| 1999 | Preserving Semantic Dependencies in Synchronous Tree Adjoining GrammarabstractRambow, Wier and Vijay-Shanker (Rambow et al., 1995) point out the differences between TAG derivation structures and semantic or predicate-argument dependencies, and Joshi and Vijay-Shanker (Joshi and Vijay-Shanker, 1999) describe a monotonic compositional semantics based on attachment order that represents the desired dependencies of a derivation without underspecifying predicate-argument relationships at any stage. In this paper, we apply the Joshi and Vijay-Shanker conception of compositional semantics to the problem of preserving semantic dependencies in Synchronous TAG translation (Shieber and Schabes, 1990; Abeillé et al., 1990). In particular, we describe an algorithm to obtain the semantic dependencies on a TAG parse forest and construct a target derivation forest with isomorphic or locally non-isomorphic dependencies in O (n7) time. William Schuler |
ACL | 1 |