VLDB 2026 Research / reviewers in the wild / expert
David A. Smith
dblp:45/3159
· DBLP profile ↗
52ranked-venue papers
11as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 7 first-author · 4 since 2021Databases, data management, data science and information retrieval · 20 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Theory of computation · 2Computer networks · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
12 papers |
Information extraction and text analysis · 70% Graph learning · 14% Language models and text generation · 9% | |
| Databases, data mining, and information retrieval
6 papers |
Information retrieval · 60% Web and social media mining · 20% Machine learning and data management · 12% | |
| Human-computer interaction and pervasive computing
2 papers |
Collaborative and social computing · 87% Immersive interaction · 13% | |
| Computer graphics and multimedia
2 papers |
Virtual and augmented reality · 100% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 67% Graph algorithms and graph theory · 33% |
Topics — the 30 heaviest of 43, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
syntactic parsing |
0.4 | 6 | 2009 | Parser Adaptation and Projection with Quasi-Synchronous Grammar Features · EMNLP 2009 Dependency Parsing by Belief Propagation · EMNLP 2008 Probabilistic Models of Nonprojective Dependency Trees · EMNLP-CoNLL 2007 |
Collaborative and social computing › mixed reality collaboration
collaborative augmented reality |
0.4 | 1 | 2020 | The Augmented Conversation and the Amplified World · UIST 2020 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
dependency parsing |
0.4 | 5 | 2012 | Parse, Price and Cut--Delayed Column and Row Generation for Graph Based Parsers · EMNLP-CoNLL 2012 Dependency Parsing by Belief Propagation · EMNLP 2008 Probabilistic Models of Nonprojective Dependency Trees · EMNLP-CoNLL 2007 |
Machine learning › Graph learning
graph structure learning |
0.3 | 1 | 2018 | Contrastive Training for Models of Information Cascades · AAAI 2018 |
Web and social media mining
information diffusion |
0.3 | 1 | 2018 | Contrastive Training for Models of Information Cascades · AAAI 2018 |
Information retrieval
evaluation |
0.2 | 1 | 2015 | Evaluating Retrieval Models through Histogram Analysis · SIGIR 2015 |
Machine learning and data management › model evaluation
retrieval model evaluation |
0.2 | 1 | 2015 | Evaluating Retrieval Models through Histogram Analysis · SIGIR 2015 |
Virtual and augmented reality › immersive interaction
collaborative virtual environments |
0.2 | 1 | 2014 | The Virtual World Framework: Collaborative virtual environments on the web · VR 2014 |
Natural language and speech › Information extraction and text analysis › syntactic parsing › dependency parsing
graph-based parsing |
0.1 | 1 | 2012 | Parse, Price and Cut--Delayed Column and Row Generation for Graph Based Parsers · EMNLP-CoNLL 2012 |
Mathematical optimization
integer programming |
0.1 | 1 | 2012 | Parse, Price and Cut--Delayed Column and Row Generation for Graph Based Parsers · EMNLP-CoNLL 2012 |
Information retrieval
query processing |
0.1 | 2 | 2011 | Two-stage query segmentation for information retrieval · SIGIR 2009 Joint Annotation of Search Queries · ACL 2011 |
Collaborative and social computing
computer-mediated communication |
0.1 | 1 | 2020 | The Augmented Conversation and the Amplified World · UIST 2020 |
Natural language and speech › Information extraction and text analysis › morphological analysis
morphological disambiguation |
0.1 | 1 | 2011 | A Discriminative Model for Joint Morphological Disambiguation and Dependency Parsing · ACL 2011 |
Information retrieval › query understanding
query annotation |
0.1 | 1 | 2011 | Joint Annotation of Search Queries · ACL 2011 |
Information retrieval
query understanding |
0.1 | 1 | 2011 | Joint Annotation of Search Queries · ACL 2011 |
Natural language and speech › Information extraction and text analysis › multilingual NLP
cross-lingual projection |
0.1 | 1 | 2009 | Parser Adaptation and Projection with Quasi-Synchronous Grammar Features · EMNLP 2009 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
parser adaptation |
0.1 | 1 | 2009 | Parser Adaptation and Projection with Quasi-Synchronous Grammar Features · EMNLP 2009 |
Natural language and speech › Information extraction and text analysis
topic model |
0.1 | 1 | 2009 | Polylingual Topic Models · EMNLP 2009 |
Information retrieval › query understanding › query parsing
query segmentation |
0.1 | 1 | 2009 | Two-stage query segmentation for information retrieval · SIGIR 2009 |
Information retrieval
retrieval models |
0.1 | 1 | 2009 | Two-stage query segmentation for information retrieval · SIGIR 2009 |
Information retrieval › retrieval models › language model
term dependency models |
0.1 | 1 | 2009 | Two-stage query segmentation for information retrieval · SIGIR 2009 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
belief propagation |
0.1 | 1 | 2008 | Dependency Parsing by Belief Propagation · EMNLP 2008 |
Natural language and speech › Information extraction and text analysis
bootstrapping |
0.1 | 1 | 2007 | Bootstrapping Feature-Rich Dependency Parsers with Entropic Priors · EMNLP-CoNLL 2007 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.1 | 1 | 2007 | Probabilistic Models of Nonprojective Dependency Trees · EMNLP-CoNLL 2007 |
Natural language and speech › Information extraction and text analysis › syntactic parsing › dependency parsing
non-projective dependency parsing |
0.1 | 1 | 2007 | Log-Linear Models of Non-Projective Trees, $k$-best MST Parsing and Tree-Ranking · EMNLP-CoNLL 2007 |
Graph algorithms and graph theory › spanning tree
minimum spanning tree |
0.1 | 1 | 2007 | Log-Linear Models of Non-Projective Trees, $k$-best MST Parsing and Tree-Ranking · EMNLP-CoNLL 2007 |
Data mining › text mining › topic modeling
latent dirichlet allocation |
0.1 | 1 | 2015 | Evaluating Retrieval Models through Histogram Analysis · SIGIR 2015 |
Data mining › text mining
topic model |
0.1 | 1 | 2015 | Evaluating Retrieval Models through Histogram Analysis · SIGIR 2015 |
Natural language and speech › Information extraction and text analysis › multilingual NLP
cross-lingual parsing |
0.0 | 1 | 2004 | Bilingual Parsing with Factored Estimation: Using English to Parse Korean · EMNLP 2004 |
Information retrieval
search engines |
0.0 | 1 | 2011 | Joint Annotation of Search Queries · ACL 2011 |
Methods — techniques the papers use, named apart from their topics
unsupervised learning · 0.7contrastive training · 0.7web standards · 0.43d multiuser framework · 0.4row generation · 0.3column generation · 0.3histogram analysis · 0.2marginalization · 0.1latent variable modeling · 0.1discriminative model · 0.1quasi-synchronous grammar · 0.1latent dirichlet allocation · 0.1belief propagation · 0.1log-linear model · 0.1k-best MST parsing · 0.1statistical significance measures · 0.0collocation analysis · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Through the Lens of History: Methods for Analyzing Temporal Variation in Content and Framing of State-run Chinese NewspapersabstractShijia Liu, David A. Smith. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Shijia Liu, David A. Smith |
NAACL (Long Papers) | 2 |
| 2024 | Self-training and Active Learning with Pseudo-relevance Feedback for Handwriting Detection in Historical Print
Jacob Murel, David A. Smith |
ICDAR (3) | 2 |
| 2024 | Detecting Manuscript Annotations in Historical Print: Negative Evidence and Evaluation Metrics
Jacob Murel, David A. Smith |
ICPRAM | 2 |
| 2021 | Content-based Models of QuotationabstractWe explore the task of quotability identification, in which, given a document, we aim to identify which of its passages are the most quotable, i.e. the most likely to be directly quoted by later derived documents. We approach quotability identification as a passage ranking problem and evaluate how well both feature-based and BERT-based (Devlin et al., 2019) models rank the passages in a given document by their predicted quotability. We explore this problem through evaluations on five datasets that span multiple languages (English, Latin) and genres of literature (e.g. poetry, plays, novels) and whose corresponding derived documents are of multiple types (news, journal articles). Our experiments confirm the relatively strong performance of BERT-based models on this task, with the best model, a RoBERTA sequential sentence tagger, achieving an average rho of 0.35 and NDCG@1, 5, 50 of 0.26, 0.31 and 0.40, respectively, across all five datasets. Ansel MacLaughlin, David A. Smith |
EACL | 2 |
| 2021 | Digital Editions as Distant Supervision for Layout Analysis of Printed Books
Alejandro H. Toselli, David A. Smith |
ICDAR (2) | 3 |
| 2020 | Detecting de minimis Code-Switching in Historical German BooksabstractCode-switching has long interested linguists, with computational work in particular focusing on speech and social media data (Sitaram et al., 2019).This paper contrasts these informal instances of code-switching to its appearance in more formal registers, by examining the mixture of languages in the Deutsches Textarchiv (DTA), a corpus of 1406 primarily German books from the 17th to 19th centuries.We automatically annotate and manually inspect spans of six embedded languages (Latin, French, English, Italian, Spanish, and Greek) in the corpus.We quantitatively analyze the differences between code-switching patterns in these books and those in more typically studied speech and social media corpora.Furthermore, we address the practical task of predicting code-switching from features of the matrix language alone in the DTA corpus.Such classifiers can help reduce errors when optical character recognition or speech transcription is applied to a large corpus with rare embedded languages. Shijia Liu, David A. Smith |
COLING | 2 |
| 2020 | Source Attribution: Recovering the Press Releases Behind Health Science News
Ansel MacLaughlin, John Wihbey, Aleszu Bajak, David A. Smith |
ICWSM | 4 |
| 2020 | The Augmented Conversation and the Amplified WorldabstractHuman communication mediated by computers and Augmented Reality devices will enable us to dynamically express, share and explore new ideas with each other via live simulations as easily as we talk about the weather. This collaboration provides a "shared truth" - what you see is exactly what I see, I see you perform an action as you do it, and we both see exactly the same dynamic transformation of this shared information space. When you express an idea, the computer, a full participant in this conversation, instantly makes it real for both of us enabling us to critique and negotiate the meaning of it. This shared virtual world will be as live, dynamic, pervasive, and visceral as the physical. David A. Smith |
UIST | 1 |
| 2018 | Contrastive Training for Models of Information CascadesabstractThis paper proposes a model of information cascades as directed spanning trees (DSTs) over observed documents. In addition, we propose a contrastive training procedure that exploits partial temporal ordering of node infections in lieu of labeled training links. This combination of model and unsupervised training makes it possible to improve on models that use infection times alone and to exploit arbitrary features of the nodes and of the text content of messages in information cascades. With only basic node and time lag features similar to previous models, the DST model achieves performance with unsupervised training comparable to strong baselines on a blog network inference task. Unsupervised training with additional content features achieves significantly better results, reaching half the accuracy of a fully supervised model. Shaobin Xu, David A. Smith |
AAAI | 2 |
| 2018 | Predicting News Coverage of Scientific Articles
Ansel MacLaughlin, John Wihbey, David A. Smith |
ICWSM | 3 |
| 2017 | A perspective from the long view: 35 Years in VR (Keynote)abstractI have been working in VR and interactive 3D for a long time. I have had the pleasure of knowing and working with many of the people that are directly responsible for creating the magic in the world that we live in today. These are the people that started with a virtual blank page and created their own reality. Their vision defined a vector into their future that we have had the privilege of extending into ours. Knowing where this vector started gives us an incredible perspective on where it is today and where it is going. I will describe my personal journey along this vector over the last 35 years, demonstrate a few things I am working on today, and speculate about where this vector into the future may take us. David A. Smith |
VR | 1 |
| 2016 | Bootstrapping Translation Detection and Sentence Extraction from Comparable Corpora
Kriste Krstovski, David A. Smith |
HLT-NAACL | 2 |
| 2016 | Online Multilingual Topic Models with Multi-Level HyperpriorsabstractFor topic models, such as LDA, that use a bag-of-words assumption, it becomes especially important to break the corpus into appropriately-sized "documents".Since the models are estimated solely from the term cooccurrences, extensive documents such as books or long journal articles lead to diffuse statistics, and short documents such as forum posts or product reviews can lead to sparsity.This paper describes practical inference procedures for hierarchical models that smooth topic estimates for smaller sections with hyperpriors over larger documents.Importantly for large collections, these online variational Bayes inference methods perform a single pass over a corpus and achieve better perplexity than "flat" topic models on monolingual and multilingual data.Furthermore, on the task of detecting document translation pairs in large multilingual collections, polylingual topic models (PLTM) with multi-level hyperpriors (mlhPLTM) achieve significantly better performance than existing online PLTM models while retaining computational efficiency. Kriste Krstovski, David A. Smith, Michael J. Kurtz |
HLT-NAACL | 2 |
| 2015 | Evaluating Retrieval Models through Histogram AnalysisabstractWe present a novel approach for efficiently evaluating the performance of retrieval models and introduce two evaluation metrics: Distributional Overlap (DO), which compares the clustering of scores of relevant and non-relevant documents, and Histogram Slope Analysis (HSA), which examines the log of the empirical distributions of relevant and non-relevant documents. Unlike rank evaluation metrics such as mean average precision (MAP) and normalized discounted cumulative gain (NDCG), DO and HSA only require calculating model scores of queries and a fixed sample of relevant and non-relevant documents rather than scoring the entire collection, even implicitly by means of an inverted index. In experimental meta-evaluations, we find that HSA achieves high correlation with MAP and NDCG on a monolingual and a cross-language document similarity task; on four ad-hoc web retrieval tasks; and on an analysis of ten TREC tasks from the past ten years. In addition, when evaluating latent Dirichlet allocation (LDA) models on document similarity tasks, HSA achieves better correlation with MAP and NCDG than perplexity, an intrinsic metric widely used with topic models. Kriste Krstovski, David A. Smith, Michael J. Kurtz |
SIGIR | 2 |
| 2014 | Social Network Signatures of Effective Online Communication
Xiaoxi Xu, Tom Murray 0001, Beverly P. Woolf, David A. Smith |
Intelligent Tutoring Systems | 4 |
| 2014 | The Virtual World Framework: Collaborative virtual environments on the webabstractSoftware distribution and installation is a logistical issue for large enterprises. Web applications are often a good solution because users can instantly receive application updates on any device without needing special permissions to install them on their hardware. Until recently, it was not possible to create 3D multiuser virtual environment-based web applications that didn't require installing a browser plugin. However, recent web standards have made it possible. We present the Virtual World Framework (VWF), a software framework for creating 3D multiuser web applications. We are using VWF to create applications for team training and collaboration. VWF can be downloaded at http://virtual.wf. Eric Burns, David Easter, Rob Chadwick, David A. Smith, Carl Rosengrant |
VR | 4 |
| 2014 | Automatic suggestion of phrasal-concept queries for literature search
Jangwon Seo, W. Bruce Croft, David A. Smith |
Inf. Process. Manag. | 4 |
| 2013 | Infectious texts: Modeling text reuse in nineteenth-century newspapersabstractTexts propagate through many social networks and provide evidence for their structure. We present efficient algorithms for detecting clusters of reused passages embedded within longer documents in large collections. We apply these techniques to analyzing the culture of reprinting in the United States before the Civil War. Without substantial copyright enforcement, stories, poems, news, and anecdotes circulated freely among newspapers, magazines, and books. From a collection of OCR'd newspapers, we extract a new corpus of reprinted texts, explore the geographic spread and network connections of different publications, and analyze the time dynamics of different genres. David A. Smith, Ryan Cordell, Elizabeth Maddock Dillon |
IEEE BigData | 1 |
| 2013 | Mining Social Deliberation in Online Communication - If You Were Me and I Were You
Xiaoxi Xu, Tom Murray 0001, Beverly P. Woolf, David A. Smith |
EDM | 4 |
| 2013 | Using a Probabilistic Syllable Model to Improve Scene Text RecognitionabstractThis paper presents a new language model for text recognition in natural images. Many existing techniques incorporate n-gram information as an additional source of information. One problem is that some n-grams are very uncommon, but will still appear in a word across a syllable boundary. These words are given a low probability under an n-gram model. To overcome this problem, we introduce a probabilistic syllable model that uses a probabilistic context-free grammar to generate recognized word labels that are consistent with syllables. In other words, labels generated by this model are pronounceable. This is important for scene text recognition where text often includes proper nouns and standard dictionary information cannot be a useful resource. We show that this language model leads to increased recognition accuracy over a big ram model and discuss the benefits over a dictionary model. Jacqueline L. Feild, Erik G. Learned-Miller, David A. Smith |
ICDAR | 3 |
| 2012 | Grammarless Parsing for Joint Inference
Jason Naradowsky, Tim Vieira, David A. Smith |
COLING | 3 |
| 2012 | Improving NLP through Marginalization of Hidden Syntactic Structure
Jason Naradowsky, Sebastian Riedel 0001, David A. Smith |
EMNLP-CoNLL | 3 |
| 2012 | Parse, Price and Cut--Delayed Column and Row Generation for Graph Based Parsers
Sebastian Riedel 0001, David A. Smith, Andrew McCallum |
EMNLP-CoNLL | 2 |
| 2012 | A framework for manipulating and searching multiple retrieval typesabstractConventional retrieval systems view documents as a unit and look at different retrieval types within a document. We introduce Proteus, a frame-work for seamlessly navigating books as dynamic collections which are defined on the fly. Proteus allows us to search various retrieval types. Navigable types include pages, books, named persons, locations, and pictures in a collection of books taken from the Internet Archive. The demonstration shows the value of multi-type browsing in dynamic collections to peruse new data. Marc-Allen Cartright, Ethem F. Can, William Dabney, Jeff Dalton 0001, Logan Giorda, Kriste Krstovski, Xiaoye Wu, Ismet Zeki Yalniz, James Allan 0001, R. Manmatha, David A. Smith |
SIGIR | 11 |
| 2011 | Joint Annotation of Search Queries
Michael Bendersky, W. Bruce Croft, David A. Smith |
ACL | 3 |
| 2011 | A Discriminative Model for Joint Morphological Disambiguation and Dependency Parsing
John Lee 0001, Jason Naradowsky, David A. Smith |
ACL | 3 |
| 2011 | Passage retrieval for incorporating global evidence in sequence labelingabstractMany forms of linguistic analysis, such as part of speech tagging, named entity recognition, and other sequence labeling tasks are performed on short spans of text and assume statistical dependence within a window of only a few tokens. We propose using passage retrieval to induce non-local dependencies in structured classification that generalizes earlier work in context aggregation for named-entity recognition. We introduce a new method for feature expansion inspired by psuedo-relevance feedback (PRF). Our results on the CoNLL 2003 task show that features from cross-document feature expansion improves NER effectiveness over previous aggregation models. Utilizing all the tokens in a sentence for query context consistently perform best on both intrinsic and extrinsic evaluations. Tagging models incorporating feature expansion outperform the leading NER system when evaluated on out of domain data, a collection of publicly available scanned books on the topic of historic Deerfield, MA. Finally, the results show that retrieval based feature expansion using an external collection of unlabeled text can result in further effectiveness improvements. Jeff Dalton 0001, James Allan 0001, David A. Smith |
CIKM | 3 |
| 2011 | Evaluating an associative browsing model for personal informationabstractRecent studies suggest that associative browsing can be beneficial for personal information access. Associative browsing is intuitive for the user and complements other methods of accessing personal information, such as keyword search. In our previous work, we proposed an associative browsing model of personal information in which users can navigate through the space of documents and concepts (e.g., person names, events, etc.). Our approach differs from other systems in that it presented a ranked list of associations by combining multiple measures of similarity, whose weights are improved based on click feedback from the user. Jin Young Kim 0005, W. Bruce Croft, David A. Smith, Anton Bakalov |
CIKM | 3 |
| 2011 | A quasi-synchronous dependence model for information retrievalabstractIncorporating syntactic features in a retrieval model has had very limited success in the past, with the exception of binary term dependencies. This paper presents a new term dependency modeling approach based on syntactic dependency parsing for both queries and documents. Our model is inspired by a quasi-synchronous stochastic process for machine translation[21]. We model four different types of relationships between syntactically dependent term pairs to perform inexact matching between documents and queries. We also propose a machine learning technique for predicting optimal parameter settings for a retrieval model incorporating syntactic relationships. The results on TREC collections show that the quasi-synchronous dependence model can improve retrieval performance and outperform a strong state-of-art sequential dependence baseline when we use predicted optimal parameters. W. Bruce Croft, David A. Smith |
CIKM | 3 |
| 2011 | Passage Reranking for Question Answering Using Syntactic Structures and Answer Types
Elif Aktolga, James Allan 0001, David A. Smith |
ECIR | 3 |
| 2011 | Learning on the fly: a font-free approach toward multilingual OCR
Andrew Kae, David A. Smith, Erik G. Learned-Miller |
Int. J. Document Anal. Recognit. | 2 |
| 2011 | Online community search using conversational structures
Jangwon Seo, W. Bruce Croft, David A. Smith |
Inf. Retr. | 3 |
| 2010 | Structural annotation of search queries using pseudo-relevance feedbackabstractMarking up queries with annotations such as part-of-speech tags, capitalization, and segmentation, is an important part of many approaches to query processing and understanding. Due to their brevity and idiosyncratic structure, search queries pose a challenge to existing annotation tools that are commonly trained on full-length documents. To address this challenge, we view the query as an explicit representation of a latent information need, which allows us to use pseudo-relevance feedback, and to leverage additional information from the document corpus, in order to improve the quality of query annotation. Michael Bendersky, W. Bruce Croft, David A. Smith |
CIKM | 3 |
| 2010 | Building a semantic representation for personal informationabstractA typical collection of personal information contains many documents and mentions many concepts (e.g., person names, events, etc.). In this environment, associative browsing between these concepts and documents can be useful as a complement for search. Previous approaches in the area of semantic desktops aimed at addressing this task. However, they were not practical because they require tedious manual annotation by the user. Jin Young Kim 0005, Anton Bakalov, David A. Smith, W. Bruce Croft |
CIKM | 3 |
| 2010 | Modeling reformulation using passage analysisabstractQuery reformulation modifies the original query with the aim of better matching the vocabulary of the relevant documents, and consequently improving ranking effectiveness. Previous techniques typically generate words and phrases related to the original query, but do not consider how these words and phrases would fit together in new queries. In this paper, we focus on an implementation of an approach that models reformulation as a distribution of queries, where each query is a variation of the original query. This approach considers a query as a basic unit and can capture important dependencies between words and phrases in the query. The implementation discussed here is based on passage analysis of the target corpus. Experiments on the TREC collection show that the proposed model for query reformulation significantly outperforms state-of-the-art methods. Xiaobing Xue, W. Bruce Croft, David A. Smith |
CIKM | 3 |
| 2010 | Relaxed Marginal Inference and its Application to Dependency Parsing
Sebastian Riedel 0001, David A. Smith |
HLT-NAACL | 2 |
| 2010 | Inference by Minimizing Size, Divergence, or their Sum
Sebastian Riedel 0001, David A. Smith, Andrew McCallum |
UAI | 2 |
| 2009 | Online community search using thread structureabstractOnline communities are valuable information sources where knowledge is accumulated by interactions between people. Search services provided by online community sites such as forums are often, however, quite poor. To address this, we investigate retrieval techniques that exploit the hierarchical thread structures in community sites. Since these structures are sometimes not explicit or accurately annotated, we use structure discovery techniques. We then make use of thread structures in retrieval experiments. Our results show that using thread structures that have been accurately annotated can lead to significant improvements in retrieval performance compared to strong baselines. Jangwon Seo, W. Bruce Croft, David A. Smith |
CIKM | 3 |
| 2009 | Polylingual Topic Models
David M. Mimno, Hanna M. Wallach, Jason Naradowsky, David A. Smith, Andrew McCallum |
EMNLP | 4 |
| 2009 | Parser Adaptation and Projection with Quasi-Synchronous Grammar Features
David A. Smith, Jason Eisner |
EMNLP | 1 |
| 2009 | Two-stage query segmentation for information retrievalabstractModeling term dependence has been shown to have a significant positive impact on retrieval. Current models, however, use sequential term dependencies, leading to an increased query latency, especially for long queries. In this paper, we examine two query segmentation models that reduce the number of dependencies. We find that two-stage segmentation based on both query syntactic structure and external information sources such as query logs, attains retrieval performance comparable to the sequential dependence model, while achieving a 50% reduction in query latency. Michael Bendersky, W. Bruce Croft, David A. Smith |
SIGIR | 3 |
| 2008 | Dependency Parsing by Belief Propagation
David A. Smith, Jason Eisner |
EMNLP | 1 |
| 2007 | Log-Linear Models of Non-Projective Trees, $k$-best MST Parsing and Tree-Ranking
Keith B. Hall, Jirí Havelka, David A. Smith |
EMNLP-CoNLL | 3 |
| 2007 | Bootstrapping Feature-Rich Dependency Parsers with Entropic Priors
David A. Smith, Jason Eisner |
EMNLP-CoNLL | 1 |
| 2007 | Probabilistic Models of Nonprojective Dependency Trees
David A. Smith, Noah A. Smith |
EMNLP-CoNLL | 1 |
| 2006 | Minimum Risk Annealing for Training Log-Linear Models
David A. Smith, Jason Eisner |
ACL | 1 |
| 2006 | Vine Parsing and Minimum Risk Reranking for Speed and Precision
Markus Dreyer, David A. Smith, Noah A. Smith |
CoNLL | 2 |
| 2004 | Bilingual Parsing with Factored Estimation: Using English to Parse Korean
David A. Smith, Noah A. Smith |
EMNLP | 1 |
| 2002 | Detecting and Browsing Events in Unstructured textabstractPreviews and overviews of large, heterogeneous information resources help users comprehend the scope of collections and focus on particular subsets of interest. For narrative docu-ments, questions of “what happened? where? and when?” are natural points of entry. Building on our earlier work at the Perseus Project with detecting terms, place names, and dates, we have exploited co-occurrences of dates and place names to detect and describe likely events in document col-lections. We compare statistical measures for determining the relative significance of various events. We have built in-terfaces that help users preview likely regions of interest for a given range of space and time by plotting the distribution and relevance of various collocations. Users can also control the amount of collocation information in each view. Once particular collocations are selected, the system can identify key phrases associated with each possible event to organize browsing of the documents themselves. David A. Smith |
SIGIR | 1 |
| 1990 | Integrated-Optic Acoustically-Tunable Filters for WDM NetworksabstractThe background needed to understand the integrated-optic collinear acoustically tunable optical filters (ATOFs) and the fabrication and performance of both simple and multielement acoustooptic tunable filters is presented. The most important sources of crosstalk between channels simultaneously selected by a single device are discussed. ATOFs have the combined virtues of narrow passband (subnanometer bandwidth), broad tuning range (hundreds of nanometers have been demonstrated), and simultaneous independent multiple-channel filtering. The theory and practice of collinear integrated optic acoustooptic filters is presented. Devices which involve higher-order integration are discussed, including multiple-state filters for enhanced sidelobe suppression and polarization-independent configurations. Various sources of interchannel crosstalk are derived and demonstrated in multiwavelength filtering experiments.> David A. Smith, Jane E. Baran, John J. Johnson, Kwok-Wai Cheung 0003 |
IEEE J. Sel. Areas Commun. | 1 |
| 1983 | HURRY: An Acceleration Algorithm for Scalar Sequences and SeriesabstractWe present a general acceleration algorithm for alternating and monotone scalar sequences and series.The main components of the algomthm are the subroutines WHIZ and HURRY.WHIZ is a recursive implementation of Levin's u transform, and HURRY {which calls WHIZ) estimates truncation and round-off errors to make a near-optimal stopping decision and provide a very good estimate of the accuracy of the computed answer.We also present a test driver program that demonstrates the capabilities of HURRY when applied to a wide varmty of convergent and divergent sequences and series Categories and Subject Descriptors: G 1 0 [Numerical Analysis]: General--numerical algortthms; G Theodore Fessler, William F. Ford, David A. Smith |
ACM Trans. Math. Softw. | 3 |
| 1983 | Algorithm 602: HURRY: An Acceleration Algorithm for Scalar Sequences and SeriesabstractThe algorithm presented here accompanies [1], which gives the description, test results, and references.The algorithm consists of three FORTRAN subroutines, HURRY(), WHIZ(), and W H I Z l ( ) .HURRY( ) controls the logical flow, makes the stopping decision, and calls both WHIZ( ) and VCHIZI(), which implement Levin's u transform for acceleration of convergence.The two implementations are identical except that W H I Z l ( ) avoids recomputation of initial information, and thus is called only after WHIZ( ) has been called once.Also included with the algorithm is a demonstration driver program called XACCEL, which is described in [1, Sect.4], and sample input data that were used to produce the output summarized in [1, Table I]. Theodore Fessler, William Ford, David A. Smith |
ACM Trans. Math. Softw. | 3 |