VLDB 2026 Research / reviewers in the wild / expert
Stuart M. Shieber
dblp:s/StuartMShieber
· DBLP profile ↗
63ranked-venue papers
18as first author
2since 2021 · last 2023
0000-0002-7733-8195ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 17 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 11 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-authorTheory of computation · 3Applied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
21 papers |
Language models and text generation · 39% Trustworthy machine learning · 28% Information extraction and text analysis · 18% | |
| Human-computer interaction and pervasive computing
4 papers |
Learning and educational technologies · 64% Games and playful interaction · 19% Collaborative and social computing · 11% |
Topics — the 30 heaviest of 74, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 2 | 2021 | Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models · ACL/IJCNLP (1) 2021 Investigating Gender Bias in Language Models Using Causal Mediation Analysis · NeurIPS 2020 |
Natural language and speech › Language models and text generation
text generation |
0.8 | 4 | 2017 | Challenges in Data-to-Document Generation · EMNLP 2017 Adapting Sequence Models for Sentence Correction · EMNLP 2017 Word Ordering Without Syntax · EMNLP 2016 |
Natural language and speech › Information extraction and text analysis › document understanding › legal text analysis
patent approval prediction |
0.7 | 1 | 2023 | The Harvard USPTO Patent Dataset: A Large-Scale, Well-Structured, and Multi-Purpose Corpus of Patent Applications · NeurIPS 2023 |
Machine learning › Trustworthy machine learning › interpretability
causal analysis |
0.5 | 1 | 2021 | Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models · ACL/IJCNLP (1) 2021 |
Machine learning › Trustworthy machine learning › fairness
bias evaluation |
0.4 | 1 | 2020 | Investigating Gender Bias in Language Models Using Causal Mediation Analysis · NeurIPS 2020 |
Machine learning › Trustworthy machine learning
fairness |
0.4 | 1 | 2020 | Investigating Gender Bias in Language Models Using Causal Mediation Analysis · NeurIPS 2020 |
Machine learning › Trustworthy machine learning › fairness › gender bias
gender bias in language models |
0.4 | 1 | 2020 | Investigating Gender Bias in Language Models Using Causal Mediation Analysis · NeurIPS 2020 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal effect estimation
mediation analysis |
0.4 | 1 | 2020 | Investigating Gender Bias in Language Models Using Causal Mediation Analysis · NeurIPS 2020 |
Natural language and speech › Language models and text generation
neural language model |
0.4 | 2 | 2021 | Word Ordering Without Syntax · EMNLP 2016 Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models · ACL/IJCNLP (1) 2021 |
Natural language and speech › Language models and text generation › natural language understanding › sentence pair modeling
natural language inference |
0.4 | 1 | 2019 | Don't Take the Premise for Granted: Mitigating Artifacts in Natural Language Inference · ACL (1) 2019 |
Machine learning › Trustworthy machine learning
robustness |
0.4 | 1 | 2019 | Don't Take the Premise for Granted: Mitigating Artifacts in Natural Language Inference · ACL (1) 2019 |
Natural language and speech › Language models and text generation
controllable text generation |
0.3 | 1 | 2018 | Learning Neural Templates for Text Generation · EMNLP 2018 |
Natural language and speech › Language models and text generation › text generation
interpretable text generation |
0.3 | 1 | 2018 | Learning Neural Templates for Text Generation · EMNLP 2018 |
Natural language and speech › Language models and text generation › text generation
template-based generation |
0.3 | 1 | 2018 | Learning Neural Templates for Text Generation · EMNLP 2018 |
Natural language and speech › Language models and text generation › text generation
data-to-text generation |
0.3 | 1 | 2017 | Challenges in Data-to-Document Generation · EMNLP 2017 |
Natural language and speech › Language models and text generation › text generation
grammatical error correction |
0.3 | 1 | 2017 | Adapting Sequence Models for Sentence Correction · EMNLP 2017 |
Natural language and speech › Machine translation › statistical machine translation
phrase-based translation |
0.3 | 1 | 2017 | Adapting Sequence Models for Sentence Correction · EMNLP 2017 |
Natural language and speech › Machine translation
statistical machine translation |
0.3 | 1 | 2017 | Adapting Sequence Models for Sentence Correction · EMNLP 2017 |
Natural language and speech › Language models and text generation › language modeling
LSTM language model |
0.2 | 1 | 2016 | Word Ordering Without Syntax · EMNLP 2016 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
word ordering |
0.2 | 1 | 2016 | Word Ordering Without Syntax · EMNLP 2016 |
Learning and educational technologies
language learning |
0.2 | 1 | 2016 | Ingenium: Engaging Novice Students with Latin Grammar · CHI 2016 |
Automata and formal languages
tree adjoining grammar |
0.2 | 2 | 2013 | A Context Free TAG Variant · ACL (1) 2013 Optimal k-arization of Synchronous Tree-Adjoining Grammar · ACL 2008 |
Natural language and speech › Information extraction and text analysis › coreference resolution
anaphoricity detection |
0.2 | 1 | 2015 | Learning Anaphoricity and Antecedent Ranking Features for Coreference Resolution · ACL (1) 2015 |
Natural language and speech › Information extraction and text analysis
coreference resolution |
0.2 | 1 | 2015 | Learning Anaphoricity and Antecedent Ranking Features for Coreference Resolution · ACL (1) 2015 |
Natural language and speech › Language models and text generation › text summarization
abstractive summarization |
0.2 | 1 | 2023 | The Harvard USPTO Patent Dataset: A Large-Scale, Well-Structured, and Multi-Purpose Corpus of Patent Applications · NeurIPS 2023 |
Natural language and speech › Language models and text generation
language modeling |
0.2 | 1 | 2023 | The Harvard USPTO Patent Dataset: A Large-Scale, Well-Structured, and Multi-Purpose Corpus of Patent Applications · NeurIPS 2023 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
plan recognition |
0.1 | 1 | 2012 | Plan recognition in exploratory domains · Artif. Intell. 2012 |
Knowledge, reasoning and agents › Multi-agent systems
multi-agent decision making |
0.1 | 1 | 2010 | Agent decision-making in open mixed networks · Artif. Intell. 2010 |
Natural language and speech › Machine translation › synchronous grammar
synchronous grammar induction |
0.1 | 1 | 2010 | Bayesian Synchronous Tree-Substitution Grammar Induction and Its Application to Sentence Compression · ACL 2010 |
Natural language and speech › Machine translation
tree-substitution grammar |
0.1 | 1 | 2010 | Bayesian Synchronous Tree-Substitution Grammar Induction and Its Application to Sentence Compression · ACL 2010 |
Methods — techniques the papers use, named apart from their topics
causal mediation analysis · 0.9natural language processing · 0.7classification model · 0.7probing · 0.5neuron analysis · 0.4attention head analysis · 0.4transfer learning · 0.4probabilistic modeling · 0.4hidden semi-markov model · 0.3formal language theory · 0.3encoder-decoder model · 0.3block-based programming · 0.2automated planning · 0.0constraint satisfaction · 0.0simulated annealing · 0.0gradient descent · 0.0optimization · 0.0heuristic comparison · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | The Harvard USPTO Patent Dataset: A Large-Scale, Well-Structured, and Multi-Purpose Corpus of Patent ApplicationsabstractInnovation is a major driver of economic and social development, and information about many kinds of innovation is embedded in semi-structured data from patents and patent applications. Though the impact and novelty of innovations expressed in patent data are difficult to measure through traditional means, machine learning offers a promising set of techniques for evaluating novelty, summarizing contributions, and embedding semantics. In this paper, we introduce the Harvard USPTO Patent Dataset (HUPD), a large-scale, well-structured, and multi-purpose corpus of English-language patent applications filed to the United States Patent and Trademark Office (USPTO) between 2004 and 2018. With more than 4.5 million patent documents, HUPD is two to three times larger than comparable corpora. Unlike other NLP patent datasets, HUPD contains the inventor-submitted versions of patent applications, not the final versions of granted patents, allowing us to study patentability at the time of filing using NLP methods for the first time. It is also novel in its inclusion of rich structured data alongside the text of patent filings: By providing each application’s metadata along with all of its text fields, HUPD enables researchers to perform new sets of NLP tasks that leverage variation in structured covariates. As a case study on the types of research HUPD makes possible, we introduce a new task to the NLP community -- patent acceptance prediction. We additionally show the structured metadata provided in HUPD allows us to conduct explicit studies of concept shifts for this task. We find that performance on patent acceptance prediction decays when models trained in one context are evaluated on different innovation categories and over time. Finally, we demonstrate how HUPD can be used for three additional tasks: Multi-class classification of patent subject areas, language modeling, and abstractive summarization. Put together, our publicly-available dataset aims to advance research extending language and classification models to diverse and dynamic real-world data distributions. Mirac Suzgun, Luke Melas-Kyriazi, Suproteem K. Sarkar, Scott Duke Kominers, Stuart M. Shieber |
NeurIPS | 5 |
| 2021 | Causal Analysis of Syntactic Agreement Mechanisms in Neural Language ModelsabstractMatthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart Shieber, Tal Linzen, Yonatan Belinkov. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart M. Shieber, Tal Linzen, Yonatan Belinkov |
ACL/IJCNLP (1) | 4 |
| 2020 | Investigating Gender Bias in Language Models Using Causal Mediation AnalysisabstractMany interpretation methods for neural models in natural language processing investigate how information is encoded inside hidden representations. However, these methods can only measure whether the information exists, not whether it is actually used by the model. We propose a methodology grounded in the theory of causal mediation analysis for interpreting which parts of a model are causally implicated in its behavior. The approach enables us to analyze the mechanisms that facilitate the flow of information from input to output through various model components, known as mediators. As a case study, we apply this methodology to analyzing gender bias in pre-trained Transformer language models. We study the role of individual neurons and attention heads in mediating gender bias across three datasets designed to gauge a model's sensitivity to gender bias. Our mediation analysis reveals that gender bias effects are concentrated in specific components of the model that may exhibit highly specialized behavior. Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, Stuart M. Shieber |
NeurIPS | 7 |
| 2019 | Don't Take the Premise for Granted: Mitigating Artifacts in Natural Language InferenceabstractNatural Language Inference (NLI) datasets often contain hypothesis-only biases-artifacts that allow models to achieve non-trivial performance without learning whether a premise entails a hypothesis.We propose two probabilistic methods to build models that are more robust to such biases and better transfer across datasets.In contrast to standard approaches to NLI, our methods predict the probability of a premise given a hypothesis and NLI label, discouraging models from ignoring the premise.We evaluate our methods on synthetic and existing NLI datasets by training on datasets containing biases and testing on datasets containing no (or different) hypothesis-only biases.Our results indicate that these methods can make NLI models more robust to dataset-specific artifacts, transferring better than a baseline architecture in 9 out of 12 NLI datasets.Additionally, we provide an extensive analysis of the interplay of our methods with known biases in NLI datasets, as well as the effects of encouraging models to ignore biases and fine-tuning on target datasets.1 * * Equal contribution 1 Our code is available at https://github.com/ azpoliak/robust-nli.2 This hypothesis contradicts the premise and would likely not be inferred. Yonatan Belinkov, Adam Poliak, Stuart M. Shieber, Benjamin Van Durme, Alexander M. Rush |
ACL (1) | 3 |
| 2019 | Automatically Analyzing Brainstorming Language Behavior with MeeterabstractLanguage both influences and indicates group behavior, and we need tools that let us study the content of what is communicated. While one could annotate these spoken dialogue acts by hand, this is a tedious, not scalable process. We present Meeter, a tool for automatically detecting information sharing, shared understanding, word counts, and group activation in spoken interactions. The contribution of our work is two-fold: (1) We validated the tool by showing that the measures computed by Meeter align with human-generated labels, and (2) we demonstrated the value of Meeter as a research tool by quantifying aspects of group behavior using those measures and deriving novel findings from that. Our tool is valuable for researchers conducting group science, as well as those designing groupware systems. Bernd Huber, Stuart M. Shieber, Krzysztof Z. Gajos |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2018 | Learning Neural Templates for Text GenerationabstractWhile neural, encoder-decoder models have had significant empirical success in text generation, there remain several unaddressed problems with this style of generation.Encoderdecoder models are largely (a) uninterpretable, and (b) difficult to control in terms of their phrasing or content.This work proposes a neural generation system using a hidden semimarkov model (HSMM) decoder, which learns latent, discrete templates jointly with learning to generate.We show that this model learns useful templates, and that these templates make generation both more interpretable and controllable.Furthermore, we show that this approach scales to real data sets and achieves strong performance nearing that of encoderdecoder text generation models. Sam Wiseman, Stuart M. Shieber, Alexander M. Rush |
EMNLP | 2 |
| 2017 | Adapting Sequence Models for Sentence CorrectionabstractIn a controlled experiment of sequence-tosequence approaches for the task of sentence correction, we find that characterbased models are generally more effective than word-based models and models that encode subword information via convolutions, and that modeling the output data as a series of diffs improves effectiveness over standard approaches.Our strongest sequence-to-sequence model improves over our strongest phrase-based statistical machine translation model, with access to the same data, by 6 M 2 (0.5 GLEU) points.Additionally, in the data environment of the standard CoNLL-2014 setup, we demonstrate that modeling (and tuning against) diffs yields similar or better M 2 scores with simpler models and/or significantly less data than previous sequence-to-sequence approaches. Allen Schmaltz, Alexander M. Rush, Stuart M. Shieber |
EMNLP | 4 |
| 2017 | Challenges in Data-to-Document GenerationabstractRecent neural models have shown significant progress on the problem of generating short descriptive texts conditioned on a small number of database records.In this work, we suggest a slightly more difficult data-to-text generation task, and investigate how effective current approaches are on this task.In particular, we introduce a new, large-scale corpus of data records paired with descriptive documents, propose a series of extractive evaluation methods for analyzing performance, and obtain baseline results using current neural generation methods.Experiments show that these models produce fluent text, but fail to convincingly approximate humangenerated documents.Moreover, even templated baselines exceed the performance of these neural models on some metrics, though copy-and reconstructionbased extensions lead to noticeable improvements. Sam Wiseman, Stuart M. Shieber, Alexander M. Rush |
EMNLP | 2 |
| 2017 | CRADLE: An Online Plan Recognition Algorithm for Exploratory DomainsabstractIn exploratory domains, agents’ behaviors include switching between activities, extraneous actions, and mistakes. Such settings are prevalent in real world applications such as interaction with open-ended software, collaborative office assistants, and integrated development environments. Despite the prevalence of such settings in the real world, there is scarce work in formalizing the connection between high-level goals and low-level behavior and inferring the former from the latter in these settings. We present a formal grammar for describing users’ activities in such domains. We describe a new top-down plan recognition algorithm called CRADLE (Cumulative Recognition of Activities and Decreasing Load of Explanations) that uses this grammar to recognize agents’ interactions in exploratory domains. We compare the performance of CRADLE with state-of-the-art plan recognition algorithms in several experimental settings consisting of real and simulated data. Our results show that CRADLE was able to output plans exponentially more quickly than the state-of-the-art without compromising its correctness, as determined by domain experts. Our approach can form the basis of future systems that use plan recognition to provide real-time support to users in a growing class of interesting and challenging domains. Reuth Mirsky, Kobi Gal, Stuart M. Shieber |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2016 | Ingenium: Engaging Novice Students with Latin GrammarabstractReading Latin poses many difficulties for English speakers, because they are accustomed to relying on word order to determine the roles of words in a sentence. In Latin, the grammatical form of a word, and not its position, is responsible for determining the word's function in a sentence. It has proven challenging to develop pedagogical techniques that successfully draw students' attention to the grammar of Latin and that students find engaging enough to use. Building on some of the most promising prior work in Latin instruction-the Michigan Latin approach--and on the insights underlying block-based programming languages used to teach children the basics of computer science, we developed Ingenium. Ingenium uses abstract puzzle blocks to communicate grammatical concepts. Engaging students in grammatical reflection, Ingenium succeeds when students are able to effectively decipher the meaning of Latin sentences. We adapted Ingenium to be used for two standard classroom activities: sentence translations and fill-in-the-blank exercises. We evaluated Ingenium with 67 novice Latin students in universities across the USA. When using Ingenium, participants opted to perform more optional exercises, completed translation exercises with significantly fewer errors related to word order and errors overall, as well as reported higher levels of engagement and attention to grammar than when using a traditional text-based interface. Sharon Zhou, Ivy J. Livingston, Mark Schiefsky, Stuart M. Shieber, Krzysztof Z. Gajos |
CHI | 4 |
| 2016 | Word Ordering Without SyntaxabstractRecent work on word ordering has argued that syntactic structure is important, or even required, for effectively recovering the order of a sentence. We find that, in fact, an n-gram language model with a simple heuristic gives strong results on this task. Furthermore, we show that a long short-term memory (LSTM) language model is even more effective at recovering order, with our basic model outperforming a state-of-the-art syntactic model by 11.5 BLEU points. Additional data and larger beams yield further gains, at the expense of training and search time. Allen Schmaltz, Alexander M. Rush, Stuart M. Shieber |
EMNLP | 3 |
| 2016 | Learning Global Features for Coreference ResolutionabstractThere is compelling evidence that coreference prediction would benefit from modeling global information about entity-clusters.Yet, state-of-the-art performance can be achieved with systems treating each mention prediction independently, which we attribute to the inherent difficulty of crafting informative clusterlevel features.We instead propose to use recurrent neural networks (RNNs) to learn latent, global representations of entity clusters directly from their mentions.We show that such representations are especially useful for the prediction of pronominal mentions, and can be incorporated into an end-to-end coreference system that outperforms the state of the art without requiring any additional search. Sam Wiseman, Alexander M. Rush, Stuart M. Shieber |
HLT-NAACL | 3 |
| 2015 | Learning Anaphoricity and Antecedent Ranking Features for Coreference ResolutionabstractSam Wiseman, Alexander M. Rush, Stuart Shieber, Jason Weston. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Sam Wiseman, Alexander M. Rush, Stuart M. Shieber, Jason Weston |
ACL (1) | 3 |
| 2014 | Eliciting and Annotating Uncertainty in Spoken Language
Heather Pon-Barry, Stuart M. Shieber, Nicholas Longenbaugh |
LREC | 2 |
| 2013 | A Context Free TAG Variant
Benjamin Swanson, Elif Yamangil, Eugene Charniak, Stuart M. Shieber |
ACL (1) | 4 |
| 2012 | Plan recognition in exploratory domains
Kobi Gal, Swapna Reddy, Stuart M. Shieber, Andee Rubin, Barbara J. Grosz |
Artif. Intell. | 3 |
| 2010 | Bayesian Synchronous Tree-Substitution Grammar Induction and Its Application to Sentence Compression
Elif Yamangil, Stuart M. Shieber |
ACL | 2 |
| 2010 | Agent decision-making in open mixed networks
Kobi Gal, Barbara J. Grosz, Sarit Kraus, Avi Pfeffer, Stuart M. Shieber |
Artif. Intell. | 5 |
| 2010 | Complexity, Parsing, and Factorization of Tree-Local Multi-Component Tree-Adjoining GrammarabstractTree-Local Multi-Component Tree-Adjoining Grammar (TL-MCTAG) is an appealing formalism for natural language representation because it arguably allows the encapsulation of the appropriate domain of locality within its elementary structures. Its multicomponent structure allows modeling of lexical items that may ultimately have elements far apart in a sentence, such as quantifiers and wh-words. When used as the base formalism for a synchronous grammar, its flexibility allows it to express both the close relationships and the divergent structure necessary to capture the links between the syntax and semantics of a single language or the syntax of two different languages. Its limited expressivity provides constraints on movement and, we posit, may have generated additional popularity based on a misconception about its parsing complexity. Although TL-MCTAG was shown to be equivalent in expressivity to TAG when it was first introduced, the complexity of TL-MCTAG is still not well understood. This article offers a thorough examination of the problem of TL-MCTAG recognition, showing that even highly restricted forms of TL-MCTAG are NP-complete to recognize. However, in spite of the provable difficulty of the recognition problem, we offer several algorithms that can substantially improve processing efficiency. First, we present a parsing algorithm that improves on the baseline parsing method and runs in polynomial time when both the fan-out and rank of the input grammar are bounded. Second, we offer an optimal, efficient algorithm for factorizing a grammar to produce a strongly equivalent TL-MCTAG grammar with the rank of the grammar minimized. Rebecca Nesson, Giorgio Satta, Stuart M. Shieber |
Comput. Linguistics | 3 |
| 2009 | Identifying uncertain words within an utterance via prosodic featuresabstractWe describe an experiment that investigates whether sub-utterance prosodic features can be used to detect uncertainty at the wordlevel. That is, given an utterance that is classified as uncertain, we want to determine which word or phrase the speaker is uncertain about. We have a corpus of utterances spoken under varying degrees of certainty. Using combinations of sub-utterance prosodic features we train models to predict the level of certainty of an utterance. On a set of utterances that were perceived to be uncertain, we compare the predictions of our models for two candidate target word segmentations: (a) one with the actual word causing uncertainty as the proposed target word, and (b) one with a control word as the proposed target word. Our best model correctly identifies the word causing the uncertainty rather than the control word 91% of the time. Heather Pon-Barry, Stuart M. Shieber |
INTERSPEECH | 2 |
| 2009 | Efficiently Parsable Extensions to Tree-Local Multicomponent TAG
Rebecca Nesson, Stuart M. Shieber |
HLT-NAACL | 2 |
| 2009 | Recognition of Users' Activities Using Constraint Satisfaction
Swapna Reddy, Kobi Gal, Stuart M. Shieber |
UMAP | 3 |
| 2008 | Optimal k-arization of Synchronous Tree-Adjoining Grammar
Rebecca Nesson, Giorgio Satta, Stuart M. Shieber |
ACL | 3 |
| 2008 | Towards Collaborative Intelligent Tutors: Automated Recognition of Users' Strategies
Kobi Gal, Elif Yamangil, Stuart M. Shieber, Andee Rubin, Barbara J. Grosz |
Intelligent Tutoring Systems | 3 |
| 2007 | Abbreviated text input using language modelingabstractWe address the problem of improving the efficiency of natural language text input under degraded conditions (for instance, on mobile computing devices or by disabled users), by taking advantage of the informational redundancy in natural language. Previous approaches to this problem have been based on the idea of prediction of the text, but these require the user to take overt action to verify or select the system's predictions. We propose taking advantage of the duality between prediction and compression. We allow the user to enter text in compressed form, in particular, using a simple stipulated abbreviation method that reduces characters by 26.4%, yet is simple enough that it can be learned easily and generated relatively fluently. We decode the abbreviated text using a statistical generative model of abbreviation, with a residual word error rate of 3.3%. The chief component of this model is an n-gram language model. Because the system's operation is completely independent from the user's, the overhead from cognitive task switching and attending to the system's actions online is eliminated, opening up the possibility that the compression-based method can achieve text input efficiency improvements where the prediction-based methods have not. We report the results of a user study evaluating this method. Stuart M. Shieber, Rani Nelken |
Nat. Lang. Eng. | 1 |
| 2006 | Practical secrecy-preserving, verifiably correct and trustworthy auctionsabstractWe present a practical system for conducting sealed-bid auctions that preserves the secrecy of the bids while providing for verifiable correctness and trustworthiness of the auction. The auctioneer must accept all bids submitted and follow the published rules of the auction. No party receives any useful information about bids before the auction closes and no bidder is able to change or repudiate her bid. Our solution uses Paillier's homomorphic encryption scheme [25] for zero knowledge proofs of correctness. Only minimal cryptographic technology is required of bidders; instead of employing complex interactive protocols or multi-party computation, the single auctioneer computes optimal auction results and publishes proofs of the results' correctness. Any party can check these proofs of correctness via publicly verifiable computations on encrypted bids. The system is illustrated through application to first-price, uniform-price and second-price auctions, including multi-item auctions. Our empirical results demonstrate the practicality of our method: auctions with hundreds of bidders are within reach of a single PC, while a modest distributed computing network can accommodate auctions with thousands of bids. David C. Parkes, Michael O. Rabin, Stuart M. Shieber, Christopher Thorpe |
ICEC | 3 |
| 2006 | Does the Turing Test Demonstrate Intelligence or Not?
Stuart M. Shieber |
AAAI | 1 |
| 2006 | Towards Robust Context-Sensitive Sentence Alignment for Monolingual Corpora
Rani Nelken, Stuart M. Shieber |
EACL | 2 |
| 2006 | Unifying Synchronous Tree Adjoining Grammars and Tree Transducers via Bimorphisms
Stuart M. Shieber |
EACL | 1 |
| 2006 | Representation in stochastic search for phylogenetic tree reconstruction
Griffin M. Weber, Lucila Ohno-Machado, Stuart M. Shieber |
J. Biomed. Informatics | 3 |
| 2003 | Writer's Aid: Using a Planner in a Collaborative Interface
Tamara Babaian, Barbara J. Grosz, Stuart M. Shieber |
IJCAI | 3 |
| 2003 | Abbreviated text inputabstractWe address the problem of improving the efficiency of natural language text input under degraded conditions (for instance, on PDAs or cell phones or by disabled users) by taking advantage of the informational redundacy in natural language. Previous approaches to this problem have been based on the idea of prediction of the text, but these require the user to take overt action to verify or select the system's predictions. We propose taking advantage of the duality between prediction and compression. We allow the user to enter text in compressed form, in particular, using a simple stipulated abbreviation method that reduces characters by about 30% yet is simple enough that it can be learned easily and generated relatively fluently. Using statistical language processing techniques, we can decode the abbreviated text with a residual word error rate of about 3%, and we expect that simple adaptive methods can improve this to about 1.5%. Because the system's operation is completely independent from the user's, the overhead from cognitive task switching and attending to the system's actions online is eliminated, opening up the possibility that the compression-based method can achieve text input efficiency improvements where the prediction-based methods have not Stuart M. Shieber, Ellie Baker |
IUI | 1 |
| 2003 | Comma Restoration Using Constituency Information
Stuart M. Shieber, Xiaopeng Tao |
HLT-NAACL | 1 |
| 2002 | The LinGO Redwoods Treebank: Motivation and Preliminary Applications
Stephan Oepen, Kristina Toutanova, Stuart M. Shieber, Christopher D. Manning, Dan Flickinger, Thorsten Brants |
COLING | 3 |
| 2002 | A writer's collaborative assistantabstractIn traditional human-computer interfaces, a human master directs a computer system as a servant, telling it not only what to do, but also how to do it. Collaborative interfaces attempt to realign the roles, making the participants collaborators in solving the person's problem. This paper describes Writer's Aid, a system that deploys AI planning techniques to enable it to serve as an author's collaborative assistant. Writer's Aid differs from previous collaborative interfaces in both the kinds of actions the system partner takes and the underlying technology it uses to do so. While an author writes a document, Writer's Aid helps in identifying and inserting citation keys and by autonomously finding and caching potentially relevant papers and their associated bibliographic information from various on-line sources. This autonomy, enabled by the use of a planning system at the core of Writer's Aid, distinguishes this system from other collaborative interfaces. The collaborative design and its division of labor result in more efficient operation: faster and easier writing on the user's part and more effective information gathering on the part of the system. Subjects in our laboratory user study found the system effective and the interface intuitive and easy to use. Tamara Babaian, Barbara J. Grosz, Stuart M. Shieber |
IUI | 3 |
| 1997 | Empirical Testing of Algorithms for Variable-Sized Label PlacementabstractWe report an empirical comparision of different heuristic techniques for variablesized point-feature label placement. This work may not be copied or reproduced in whole or in part for any commercial purpose. Permission to copy in whole or in part without payment of fee is granted for nonprofit educational and research purposes provided that all such whole or partial copies include the following: a notice that such copying is by permission of Mitsubishi Electric Information Technology Center America; an acknowledgment of the authors and individual contributions to the work; and all applicable portions of the copyright notice. Copying, reproduction, or republishing for any other purpose shall require a license with payment of fee to Mitsubishi Electric Information Technology Center America. All rights reserved. Copyright c fl Mitsubishi Electric Information Technology Center America, 1997 201 Broadway, Cambridge, Massachusetts 02139 1. First printing, TR97-13, October 1997. Al... Jon Christensen, Stacy Friedman, Joe Marks, Stuart M. Shieber |
SCG | 4 |
| 1997 | Design Gallery Browsers Based on 2D and 3D Graph Drawing
Brad Andalman, Kathy Ryall, Wheeler Ruml, Joe Marks, Stuart M. Shieber |
GD | 5 |
| 1997 | Design galleries: a general approach to setting parameters for computer graphics and animationabstractArticle Design galleries: a general approach to setting parameters for computer graphics and animation Share on Authors: J. Marks MERL - A Mitsubishi Electric Research Laboratory, 201 Broadway, Cambridge, MA MERL - A Mitsubishi Electric Research Laboratory, 201 Broadway, Cambridge, MAView Profile , B. Andalman Harvard Univ. Harvard Univ.View Profile , P. A. Beardsley MERL - A Mitsubishi Electric Research Laboratory, 201 Broadway, Cambridge, MA MERL - A Mitsubishi Electric Research Laboratory, 201 Broadway, Cambridge, MAView Profile , W. Freeman MERL - A Mitsubishi Electric Research Laboratory, 201 Broadway, Cambridge, MA MERL - A Mitsubishi Electric Research Laboratory, 201 Broadway, Cambridge, MAView Profile , S. Gibson MERL - A Mitsubishi Electric Research Laboratory, 201 Broadway, Cambridge, MA MERL - A Mitsubishi Electric Research Laboratory, 201 Broadway, Cambridge, MAView Profile , J. Hodgins Georgia Tech. Georgia Tech.View Profile , T. Kang CMU CMUView Profile , B. Mirtich MERL - A Mitsubishi Electric Research Laboratory, 201 Broadway, Cambridge, MA MERL - A Mitsubishi Electric Research Laboratory, 201 Broadway, Cambridge, MAView Profile , H. Pfister MERL - A Mitsubishi Electric Research Laboratory, 201 Broadway, Cambridge, MA MERL - A Mitsubishi Electric Research Laboratory, 201 Broadway, Cambridge, MAView Profile , W. Ruml MERL - A Mitsubishi Electric Research Laboratory, 201 Broadway, Cambridge, MA MERL - A Mitsubishi Electric Research Laboratory, 201 Broadway, Cambridge, MAView Profile , K. Ryall Harvard Univ. Harvard Univ.View Profile , J. Seims Univ. of Washington Univ. of WashingtonView Profile , S. Shieber Harvard Univ. Harvard Univ.View Profile Authors Info & Claims SIGGRAPH '97: Proceedings of the 24th annual conference on Computer graphics and interactive techniquesAugust 1997 Pages 389–400https://doi.org/10.1145/258734.258887Online:03 August 1997Publication History 346citation2,992DownloadsMetricsTotal Citations346Total Downloads2,992Last 12 Months156Last 6 weeks16 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Joe Marks, Brad Andalman, Paul A. Beardsley, William T. Freeman, Sarah F. Frisken, Jessica K. Hodgins, T. Kang, Brian Mirtich, Hanspeter Pfister, Wheeler Ruml, Kathy Ryall, Joshua E. Seims, Stuart M. Shieber |
SIGGRAPH | 13 |
| 1997 | An Interactive Constraint-Based System for Drawing GraphsabstractThe GLIDE system is an interactive constraint-based editor for drawing small-and medium-sized graphs (50 nodes or fewer) that organizes the interaction in a more collaborative manner than in previous systems.Its distinguishing features are a vocabulary of specialized constraints for graph drawing, and a simple constraintsatisfaction mechanism that allows the user to manipulate the drawing while the constraints are active.These features result in a graph-drawing editor that is superior in many ways to those based on more general and powerful constraint-satisfaction methods. Kathy Ryall, Joe Marks, Stuart M. Shieber |
ACM Symposium on User Interface Software and Technology | 3 |
| 1996 | An Interactive System for Drawing Graphs
Kathy Ryall, Joe Marks, Stuart M. Shieber |
GD | 3 |
| 1996 | A Viewer for PostScript DocumentsabstractNo abstract available. Adam Ginsburg, Joe Marks, Stuart M. Shieber |
ACM Symposium on User Interface Software and Technology | 3 |
| 1995 | Semi-automatic delineation of regions in floor plansabstractWe propose a technique that uses a proximity metric for delineating partially or fully bounded regions of a scanned bitmap that depicts a building floor plan. A proximity field is defined over the bitmap, which is used both to identify the centers of subjective regions in the image and to assign pixels to regions by proximity. The region boundaries generated by the method tend to match well the subjective boundaries of regions in the image. We discuss incorporation of the technique in a semi-automated interactive system for region identification in floor plans. In contrast to area-filling techniques for delineating areal regions of images, our approach works robustly for partially bounded regions. Furthermore, the frailties of the method that do remain, unlike those of alternative techniques, are well-moderated by simple human intervention. Kathy Ryall, Stuart M. Shieber, Joe Marks, Murray Mazer |
ICDAR | 2 |
| 1995 | An Empirical Study of Algorithms for Point-Feature Label PlacementabstractA major factor affecting the clarity of graphical displays that include text labels is the degree to which labels obscure display features (including other labels) as a result of spatial overlap. Point-feature label placement (PFLP) is the problem of placing text labels adjacent to point features on a map or diagram so as to maximize legibility. This problem occurs frequently in the production of many types of informational graphics, though it arises most often in automated cartography. In this paper we present a comprehensive treatment of the PFLP problem, viewed as a type of combinatorial optimization problem. Complexity analysis reveals that the basic PFLP problem and most interesting variants of it are NP-hard. These negative results help inform a survey of previously reported algorithms for PFLP; not surprisingly, all such algorithms either have exponential time complexity or are incomplete. To solve the PFLP problem in practice, then, we must rely on good heuristic methods. We propose two new methods, one based on a discrete form of gradient descent, the other on simulated annealing, and report on a series of empirical tests comparing these and the other known algorithms for the problem. Based on this study, the first to be conducted, we identify the best approaches as a function of available computation time. Jon Christensen, Joe Marks, Stuart M. Shieber |
ACM Trans. Graph. | 3 |
| 1994 | Optimization - an emerging tool in computer graphicsabstractNo abstract available. Joe Marks, Michael F. Cohen, J. Thomas Ngo, Stuart M. Shieber, John Snyder |
SIGGRAPH | 4 |
| 1994 | Restricting the Weak-Generative Capacity of Synchronous Tree-Adjoining Grammars abstractThe formalism of synchronous tree‐adjoining grammars, a variant of standard tree‐adjoining grammars (TAG), was intended to allow the use of TAGs for language transduction in addition to language specification. In previous work, the definition of the transduction relation defined by a synchronous TAG was given by appeal to an iterative rewriting process. The rewriting definition of derivation is problematic in that it greatly extends the expressivity of the formalism and makes the design of parsing algorithms difficult if not impossible. We introduce a simple, natural definition of synchronous tree‐adjoining derivation, based on isomorphisms between standard tree‐adjoining derivations, that avoids the expressivity and implementability problems of the original rewriting definition. The decrease in expressivity, which would otherwise make the method unusable, is offset by the incorporation of an alternative definition of standard tree‐adjoining derivation, previously proposed for completely separate reasons, thereby making it practical to entertain using the natural definition of synchronous derivation. Nonetheless, some remaining problematic cases call for yel more flexibility in the definition; the isomorphism requirement may have to be relaxed. It remains for future research to rune the exact requirements on the allowable mappings. Stuart M. Shieber |
Comput. Intell. | 1 |
| 1994 | An Alternative Conception of Tree-Adjoining Derivation
Yves Schabes, Stuart M. Shieber |
Comput. Linguistics | 2 |
| 1994 | Automating the layout of network diagrams with specified visual organizationabstractNetwork diagrams are a familiar graphic form that can express many different kinds of information. The problem of automating network-diagram layout has therefore received much attention. Previous research on network-diagram layout has focused on the problem of aesthetically optimal layout, using such criteria as the number of link crossings, the sum of all link lengths, and total diagram area. In this paper the authors propose a restatement of the network-diagram layout problem in which layout-aesthetic concerns are subordinated to perceptual-organization concerns. The authors present a notation for describing the visual organization of a network diagram. This notation is used in reformulating the layout task as a constrained-optimization problem in which constraints are derived from a visual-organization specification and optimality criteria are derived from layout-aesthetic considerations. Two new heuristic algorithms are presented for this version of the layout problem: one algorithm uses a rule-based strategy for computing a layout; the other is a massively parallel genetic algorithm. The authors demonstrate the capabilities of the two algorithms by testing them on a variety of network-diagram layout problems.> Corey Kosak, Joe Marks, Stuart M. Shieber |
IEEE Trans. Syst. Man Cybern. | 3 |
| 1993 | The Problem of Logical-Form Equivalence
Stuart M. Shieber |
Comput. Linguistics | 1 |
| 1992 | An Alternative Conception of Tree-Adjoining DerivationabstractThe precise formulation of derivation for tree-adjoining grammars has important ramifications for a wide variety of uses of the formalism, from syntactic analysis to semantic interpretation and statistical language modeling. We argue that the definition of tree-adjoining derivation must be reformulated in order to manifest the proper linguistic dependencies in derivations. The particular proposal is both precisely characterizable, through a compilation to linear indexed grammars, and computationally operational, by virtue of an efficient algorithm for recognition and parsing. Yves Schabes, Stuart M. Shieber |
ACL | 2 |
| 1990 | Synchronous Tree-Adjoining Grammars
Stuart M. Shieber, Yves Schabes |
COLING | 1 |
| 1990 | Generation and Synchronous Tree-Adjoining Grammars
Stuart M. Shieber, Yves Schabes |
INLG | 1 |
| 1990 | Semantic-Head-Driven Generation
Stuart M. Shieber, Gertjan van Noord, Fernando Pereira 0003, Robert C. Moore |
Comput. Linguistics | 1 |
| 1989 | A Semantic-Head-Driven Generation Algorithm for Unification-Based FormalismsabstractWe present an algorithm for generating strings from logical form encodings that improves upon previous algorithms in that it places fewer restrictions on the class of grammars to which it is applicable. In particular, unlike an Earley deduction generator (Shieber, 1988), it allows use of semantically nonmonotonic grammars, yet unlike topdown methods, it also permits left-recursion. The enabling design feature of the algorithm is its implicit traversal of the analysis tree for the string being generated in a semantic-head-driven fashion. Stuart M. Shieber, Gertjan van Noord, Robert C. Moore, Fernando Pereira 0003 |
ACL | 1 |
| 1988 | A uniform architecture for parsing and generation
Stuart M. Shieber |
COLING | 1 |
| 1987 | An Algorithm for Generating Quantifier Scopings
Jerry R. Hobbs, Stuart M. Shieber |
Comput. Linguistics | 2 |
| 1986 | A Simple Reconstruction of GPSG
Stuart M. Shieber |
COLING | 1 |
| 1985 | Using Restriction to Extend Parsing Algorithms for Complex-Feature-Based FormalismsabstractGrammar formalisms based on the encoding of grammatical information in complex-valued feature systems enjoy some currency both in linguistics and natural-language-processing research. Such formalisms can be thought of by analogy to context-free grammars as generalizing the notion of nonterminal symbol from a finite domain of atomic elements to a possibly infinite domain of directed graph structures of a certain sort. Unfortunately, in moving to an infinite nonterminal domain, standard methods of parsing may no longer be applicable to the formalism. Typically, the problem manifests itself as gross inefficiency or even nontermination of the algorithms. In this paper, we discuss a solution to the problem of extending parsing algorithms to formalisms with possibly infinite nonterminal domains, a solution based on a general technique we call restriction. As a particular example of such an extension, we present a complete, correct, terminating extension of Earley's algorithm that uses restriction to perform top-down filtering. Our implementation of this algorithm demonstrates the drastic elimination of chart edges that can be achieved by this technique. Finally, we describe further uses for the technique---including parsing other grammar formalisms, including definite-clause grammars; extending other parsing algorithms, including LR methods and syntactic preference modeling algorithms; and efficient indexing. Stuart M. Shieber |
ACL | 1 |
| 1984 | The Semantics of Grammar Formalisms Seen as Computer LanguagesabstractThe design, implementation, and use of grammar formalisms for natural language have constituted a major branch of computational linguistics throughout its development. By viewing grammar formalisms as just a special case of computer languages, we can take advantage of the machinery of denotational semantics to provide a precise specification of their meaning. Using Dana Scott's domain theory, we elucidate the nature of the feature systems used in augmented phrase-structure grammar formalisms, in particular those of recent versions of generalized phrase structure grammar, lexical functional grammar and PATR-II, and provide a denotational semantics for a simple grammar formalism. We find that the mathematical structures developed for this purpose contain an operation of feature generalization, not available in those grammar formalisms, that can be used to give a partial account of the effect of coordination on syntactic features. Fernando Pereira 0003, Stuart M. Shieber |
COLING | 2 |
| 1984 | The Design of a Computer Language for Linguistic InformationabstractA considerable body of accumulated knowledge about the design of languages for communicating information to computers has been derived from the subfields of programming language design and semantics. It has been the goal of the PATR group at SRI to utilize a relevant portion of this knowledge in implementing tools to facilitate communication of linguistic information to computers. The PATR-II formalism is our current computer language for encoding linguistic information. This paper, a brief overview of that formalism, attempts to explicate our design decisions in terms of a set of properties that effective computer languages should incorporate. Stuart M. Shieber |
COLING | 1 |
| 1983 | Sentence Disambiguation by a Shift-Reduce Parsing TechniqueabstractNative speakers of English show definite and consistent preferences for certain readings of syntactically ambiguous sentences. A user of a natural-language-processing system would naturally expect it to reflect the same preferences. Thus, such systems must model in some way the linguistic performance as well as the linguistic competence of the native speaker. We have developed a parsing algorithm---a variant of the LALR(1) shift-reduce algorithm---that models the preference behavior of native speakers for a range of syntactic preference phenomena reported in the psycholinguistic literature, including the recent data on lexical preferences. The algorithm yields the preferred parse deterministically, without building multiple parse trees and choosing among them. As a side effect, it displays appropriate behavior in processing the much discussed garden-path sentences. The parsing algorithm has been implemented and has confirmed the feasibility of our approach to the modeling of these phenomena. Stuart M. Shieber |
ACL | 1 |
| 1983 | Formal Constraints on MetarulesabstractMetagrammatical formalisms that combine context-free phrase structure rules and metarules (MPS grammars) allow concise statement of generalizations about the syntax of natural languages. Unconstrained MPS grammars, unfortunately, are not computationally "safe." We evaluate several proposals for constraining them, basing our assessment on computational tractability and explanatory adequacy. We show that none of them satisfies both criteria, and suggest new directions for research on alternative metagrammatical formalisms. Stuart M. Shieber, Swan U. Stucky, Hans Uszkoreit, Jane J. Robinson |
ACL | 1 |
| 1983 | Sentence Disambiguation by a Shift-Reduce Parsing Technique
Stuart M. Shieber |
IJCAI | 1 |
| 1982 | Translating English into Logical FormabstractA scheme for syntax-directed translation that mirrors compositional model-theoretic semantics is discussed.The scheme is the basis for an English translation system called PArR and was used to specify a semantically interesting fragment of English, including such constructs as tense, aspect, modals, and various iexically controlled verb complement structures.PATR was embedded in a question-answering system that replied appropriately to questions requiring the computation of logical entailments. Stanley J. Rosenschein, Stuart M. Shieber |
ACL | 2 |