EDBT 2026 Demo / reviewers in the wild / expert
David A. Smith
dblp:45/3159
· DBLP profile ↗
20ranked-venue papers in the field
2as first author
2since 2021 · last 2024
—ORCID · conflict
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 16 (1 first)Other / Interdisciplinary · 3Big Data, Cloud & Distributed Data Systems · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Self-training and Active Learning with Pseudo-relevance Feedback for Handwriting Detection in Historical Print
Jacob Murel, David A. Smith |
ICDAR (3) | 2 |
| 2021 | Digital Editions as Distant Supervision for Layout Analysis of Printed Books
Alejandro H. Toselli, David A. Smith |
ICDAR (2) | 3 |
| 2020 | Source Attribution: Recovering the Press Releases Behind Health Science News
Ansel MacLaughlin, John Wihbey, Aleszu Bajak, David A. Smith |
ICWSM | 4 |
| 2018 | Predicting News Coverage of Scientific Articles
Ansel MacLaughlin, John Wihbey, David A. Smith |
ICWSM | 3 |
| 2015 | Evaluating Retrieval Models through Histogram AnalysisabstractWe present a novel approach for efficiently evaluating the performance of retrieval models and introduce two evaluation metrics: Distributional Overlap (DO), which compares the clustering of scores of relevant and non-relevant documents, and Histogram Slope Analysis (HSA), which examines the log of the empirical distributions of relevant and non-relevant documents. Unlike rank evaluation metrics such as mean average precision (MAP) and normalized discounted cumulative gain (NDCG), DO and HSA only require calculating model scores of queries and a fixed sample of relevant and non-relevant documents rather than scoring the entire collection, even implicitly by means of an inverted index. In experimental meta-evaluations, we find that HSA achieves high correlation with MAP and NDCG on a monolingual and a cross-language document similarity task; on four ad-hoc web retrieval tasks; and on an analysis of ten TREC tasks from the past ten years. In addition, when evaluating latent Dirichlet allocation (LDA) models on document similarity tasks, HSA achieves better correlation with MAP and NCDG than perplexity, an intrinsic metric widely used with topic models. Kriste Krstovski, David A. Smith, Michael J. Kurtz |
SIGIR | 2 |
| 2014 | Automatic suggestion of phrasal-concept queries for literature search
Jangwon Seo, W. Bruce Croft, David A. Smith |
Inf. Process. Manag. | 4 |
| 2013 | Infectious texts: Modeling text reuse in nineteenth-century newspapersabstractTexts propagate through many social networks and provide evidence for their structure. We present efficient algorithms for detecting clusters of reused passages embedded within longer documents in large collections. We apply these techniques to analyzing the culture of reprinting in the United States before the Civil War. Without substantial copyright enforcement, stories, poems, news, and anecdotes circulated freely among newspapers, magazines, and books. From a collection of OCR'd newspapers, we extract a new corpus of reprinted texts, explore the geographic spread and network connections of different publications, and analyze the time dynamics of different genres. David A. Smith, Ryan Cordell, Elizabeth Maddock Dillon |
IEEE BigData | 1 |
| 2013 | Using a Probabilistic Syllable Model to Improve Scene Text RecognitionabstractThis paper presents a new language model for text recognition in natural images. Many existing techniques incorporate n-gram information as an additional source of information. One problem is that some n-grams are very uncommon, but will still appear in a word across a syllable boundary. These words are given a low probability under an n-gram model. To overcome this problem, we introduce a probabilistic syllable model that uses a probabilistic context-free grammar to generate recognized word labels that are consistent with syllables. In other words, labels generated by this model are pronounceable. This is important for scene text recognition where text often includes proper nouns and standard dictionary information cannot be a useful resource. We show that this language model leads to increased recognition accuracy over a big ram model and discuss the benefits over a dictionary model. Jacqueline L. Feild, Erik G. Learned-Miller, David A. Smith |
ICDAR | 3 |
| 2012 | A framework for manipulating and searching multiple retrieval typesabstractConventional retrieval systems view documents as a unit and look at different retrieval types within a document. We introduce Proteus, a frame-work for seamlessly navigating books as dynamic collections which are defined on the fly. Proteus allows us to search various retrieval types. Navigable types include pages, books, named persons, locations, and pictures in a collection of books taken from the Internet Archive. The demonstration shows the value of multi-type browsing in dynamic collections to peruse new data. Marc-Allen Cartright, Ethem F. Can, William Dabney, Jeff Dalton 0001, Logan Giorda, Kriste Krstovski, Xiaoye Wu, Ismet Zeki Yalniz, James Allan 0001, R. Manmatha, David A. Smith |
SIGIR | 11 |
| 2011 | Passage retrieval for incorporating global evidence in sequence labelingabstractMany forms of linguistic analysis, such as part of speech tagging, named entity recognition, and other sequence labeling tasks are performed on short spans of text and assume statistical dependence within a window of only a few tokens. We propose using passage retrieval to induce non-local dependencies in structured classification that generalizes earlier work in context aggregation for named-entity recognition. We introduce a new method for feature expansion inspired by psuedo-relevance feedback (PRF). Our results on the CoNLL 2003 task show that features from cross-document feature expansion improves NER effectiveness over previous aggregation models. Utilizing all the tokens in a sentence for query context consistently perform best on both intrinsic and extrinsic evaluations. Tagging models incorporating feature expansion outperform the leading NER system when evaluated on out of domain data, a collection of publicly available scanned books on the topic of historic Deerfield, MA. Finally, the results show that retrieval based feature expansion using an external collection of unlabeled text can result in further effectiveness improvements. Jeff Dalton 0001, James Allan 0001, David A. Smith |
CIKM | 3 |
| 2011 | Evaluating an associative browsing model for personal informationabstractRecent studies suggest that associative browsing can be beneficial for personal information access. Associative browsing is intuitive for the user and complements other methods of accessing personal information, such as keyword search. In our previous work, we proposed an associative browsing model of personal information in which users can navigate through the space of documents and concepts (e.g., person names, events, etc.). Our approach differs from other systems in that it presented a ranked list of associations by combining multiple measures of similarity, whose weights are improved based on click feedback from the user. Jin Young Kim 0005, W. Bruce Croft, David A. Smith, Anton Bakalov |
CIKM | 3 |
| 2011 | A quasi-synchronous dependence model for information retrievalabstractIncorporating syntactic features in a retrieval model has had very limited success in the past, with the exception of binary term dependencies. This paper presents a new term dependency modeling approach based on syntactic dependency parsing for both queries and documents. Our model is inspired by a quasi-synchronous stochastic process for machine translation[21]. We model four different types of relationships between syntactically dependent term pairs to perform inexact matching between documents and queries. We also propose a machine learning technique for predicting optimal parameter settings for a retrieval model incorporating syntactic relationships. The results on TREC collections show that the quasi-synchronous dependence model can improve retrieval performance and outperform a strong state-of-art sequential dependence baseline when we use predicted optimal parameters. W. Bruce Croft, David A. Smith |
CIKM | 3 |
| 2011 | Passage Reranking for Question Answering Using Syntactic Structures and Answer Types
Elif Aktolga, James Allan 0001, David A. Smith |
ECIR | 3 |
| 2011 | Online community search using conversational structures
Jangwon Seo, W. Bruce Croft, David A. Smith |
Inf. Retr. | 3 |
| 2010 | Structural annotation of search queries using pseudo-relevance feedbackabstractMarking up queries with annotations such as part-of-speech tags, capitalization, and segmentation, is an important part of many approaches to query processing and understanding. Due to their brevity and idiosyncratic structure, search queries pose a challenge to existing annotation tools that are commonly trained on full-length documents. To address this challenge, we view the query as an explicit representation of a latent information need, which allows us to use pseudo-relevance feedback, and to leverage additional information from the document corpus, in order to improve the quality of query annotation. Michael Bendersky, W. Bruce Croft, David A. Smith |
CIKM | 3 |
| 2010 | Building a semantic representation for personal informationabstractA typical collection of personal information contains many documents and mentions many concepts (e.g., person names, events, etc.). In this environment, associative browsing between these concepts and documents can be useful as a complement for search. Previous approaches in the area of semantic desktops aimed at addressing this task. However, they were not practical because they require tedious manual annotation by the user. Jin Young Kim 0005, Anton Bakalov, David A. Smith, W. Bruce Croft |
CIKM | 3 |
| 2010 | Modeling reformulation using passage analysisabstractQuery reformulation modifies the original query with the aim of better matching the vocabulary of the relevant documents, and consequently improving ranking effectiveness. Previous techniques typically generate words and phrases related to the original query, but do not consider how these words and phrases would fit together in new queries. In this paper, we focus on an implementation of an approach that models reformulation as a distribution of queries, where each query is a variation of the original query. This approach considers a query as a basic unit and can capture important dependencies between words and phrases in the query. The implementation discussed here is based on passage analysis of the target corpus. Experiments on the TREC collection show that the proposed model for query reformulation significantly outperforms state-of-the-art methods. Xiaobing Xue, W. Bruce Croft, David A. Smith |
CIKM | 3 |
| 2009 | Online community search using thread structureabstractOnline communities are valuable information sources where knowledge is accumulated by interactions between people. Search services provided by online community sites such as forums are often, however, quite poor. To address this, we investigate retrieval techniques that exploit the hierarchical thread structures in community sites. Since these structures are sometimes not explicit or accurately annotated, we use structure discovery techniques. We then make use of thread structures in retrieval experiments. Our results show that using thread structures that have been accurately annotated can lead to significant improvements in retrieval performance compared to strong baselines. Jangwon Seo, W. Bruce Croft, David A. Smith |
CIKM | 3 |
| 2009 | Two-stage query segmentation for information retrievalabstractModeling term dependence has been shown to have a significant positive impact on retrieval. Current models, however, use sequential term dependencies, leading to an increased query latency, especially for long queries. In this paper, we examine two query segmentation models that reduce the number of dependencies. We find that two-stage segmentation based on both query syntactic structure and external information sources such as query logs, attains retrieval performance comparable to the sequential dependence model, while achieving a 50% reduction in query latency. Michael Bendersky, W. Bruce Croft, David A. Smith |
SIGIR | 3 |
| 2002 | Detecting and Browsing Events in Unstructured textabstractPreviews and overviews of large, heterogeneous information resources help users comprehend the scope of collections and focus on particular subsets of interest. For narrative docu-ments, questions of “what happened? where? and when?” are natural points of entry. Building on our earlier work at the Perseus Project with detecting terms, place names, and dates, we have exploited co-occurrences of dates and place names to detect and describe likely events in document col-lections. We compare statistical measures for determining the relative significance of various events. We have built in-terfaces that help users preview likely regions of interest for a given range of space and time by plotting the distribution and relevance of various collocations. Users can also control the amount of collocation information in each view. Once particular collocations are selected, the system can identify key phrases associated with each possible event to organize browsing of the documents themselves. David A. Smith |
SIGIR | 1 |