VLDB 2026 Research / reviewers in the wild / expert
Philip Resnik
dblp:p/PhilipResnik
· DBLP profile ↗
79ranked-venue papers
13as first author
9since 2021 · last 2025
0000-0002-6130-8602ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 70 · 12 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 5Databases, data management, data science and information retrieval · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ProxAnn: Use-Oriented Evaluations of Topic Models and Document ClusteringabstractAlexander Miserlis Hoyle, Lorena Calvo-Bartolomé, Jordan Lee Boyd-Graber, Philip Resnik. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Alexander Miserlis Hoyle, Lorena Calvo-Bartolomé, Jordan L. Boyd-Graber, Philip Resnik |
ACL (1) | 4 |
| 2025 | Understanding Common Ground Misalignment in Goal-Oriented Dialog: A Case-Study with Ubuntu Chat LogsabstractRupak Sarkar, Neha Srikanth, Taylor Pellegrin, Rachel Rudinger, Claire Bonial, Philip Resnik. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Rupak Sarkar, Neha Srikanth, Taylor Pellegrin, Rachel Rudinger, Claire Bonial, Philip Resnik |
ACL (1) | 6 |
| 2025 | Multimodal Biomarkers for Schizophrenia: Towards Individual Symptom Severity EstimationabstractStudies on schizophrenia assessments using deep learning typically treat it as a classification task to detect the presence or absence of the disorder, oversimplifying the condition and reducing its clinical applicability. This traditional approach overlooks the complexity of schizophrenia, limiting its practical value in healthcare settings. This study shifts the focus to individual symptom severity estimation using a multimodal approach that integrates speech, video, and text inputs. We develop unimodal models for each modality and a multimodal framework to improve accuracy and robustness. By capturing a more detailed symptom profile, this approach can help in enhancing diagnostic precision and support personalized treatment, offering a scalable and objective tool for mental health assessment. Gowtham Premananth, Philip Resnik, Sonia Bansal, Deanna L. Kelly, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2024 | A Multimodal Framework for the Assessment of the Schizophrenia Spectrum
Gowtham Premananth, Yashish M. Siriwardena, Philip Resnik, Sonia Bansal, Deanna L. Kelly, Carol Y. Espy-Wilson |
INTERSPEECH | 3 |
| 2024 | TopicGPT: A Prompt-based Topic Modeling FrameworkabstractChau Minh Pham, Alexander Hoyle, Simeng Sun, Philip Resnik, Mohit Iyyer. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Alexander Miserlis Hoyle, Simeng Sun, Philip Resnik, Mohit Iyyer |
NAACL-HLT | 4 |
| 2023 | Natural Language Decompositions of Implicit Content Enable Better Text RepresentationsabstractWhen people interpret text, they rely on inferences that go beyond the observed language itself.Inspired by this observation, we introduce a method for the analysis of text that takes implicitly communicated content explicitly into account.We use a large language model to produce sets of propositions that are inferentially related to the text that has been observed, then validate the plausibility of the generated content via human judgments.Incorporating these explicit representations of implicit content proves useful in multiple problem settings that involve the human interpretation of utterances: assessing the similarity of arguments, making sense of a body of opinion data, and modeling legislative behavior.Our results suggest that modeling the meanings behind observed language, rather than the literal text alone, is a valuable direction for NLP and particularly its applications to social science. 1 * Equal contribution. Alexander Miserlis Hoyle, Rupak Sarkar, Pranav Goel 0001, Philip Resnik |
EMNLP | 4 |
| 2022 | Bernice: A Multilingual Pre-trained Encoder for TwitterabstractThe language of Twitter differs significantly from that of other domains commonly included in large language model training.While tweets are typically multilingual and contain informal language, including emoji and hashtags, most pre-trained language models for Twitter are either monolingual, adapted from other domains rather than trained exclusively on Twitter, or are trained on a limited amount of in-domain Twitter data.We introduce Bernice, the first multilingual RoBERTa language model trained from scratch on 2.5 billion tweets with a custom tweet-focused tokenizer.We evaluate on a variety of monolingual and multilingual Twitter benchmarks, finding that our model consistently exceeds or matches the performance of a variety of models adapted to social media data as well as strong multilingual baselines, despite being trained on less data overall.We posit that it is more efficient compute-and data-wise to train completely on in-domain data with a specialized domain-specific tokenizer. Alexandra DeLucia, Aaron Mueller, Carlos Alejandro Aguirre, Philip Resnik, Mark Dredze |
EMNLP | 5 |
| 2021 | Syntopical Graphs for Computational Argumentation TasksabstractJoe Barrow, Rajiv Jain, Nedim Lipka, Franck Dernoncourt, Vlad Morariu, Varun Manjunatha, Douglas Oard, Philip Resnik, Henning Wachsmuth. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Joe Barrow, Rajiv Jain, Nedim Lipka, Franck Dernoncourt, Vlad I. Morariu, Varun Manjunatha, Douglas W. Oard, Philip Resnik, Henning Wachsmuth |
ACL/IJCNLP (1) | 8 |
| 2021 | Is Automated Topic Model Evaluation Broken? The Incoherence of CoherenceabstractTopic model evaluation, like evaluation of other unsupervised methods, can be contentious. However, the field has coalesced around automated estimates of topic coherence, which rely on the frequency of word co-occurrences in a reference corpus. Contemporary neural topic models surpass classical ones according to these metrics. At the same time, topic model evaluation suffers from a validation gap: automated coherence, developed for classical models, has not been validated using human experimentation for neural models. In addition, a meta-analysis of topic modeling literature reveals a substantial standardization gap in automated topic modeling benchmarks. To address the validation gap, we compare automated coherence with the two most widely accepted human judgment tasks: topic rating and word intrusion. To address the standardization gap, we systematically evaluate a dominant classical model and two state-of-the-art neural models on two commonly used datasets. Automated evaluations declare a winning model when corresponding human evaluations do not, calling into question the validity of fully automatic evaluations independent of human judgments. Alexander Miserlis Hoyle, Pranav Goel 0001, Andrew Hian-Cheong, Denis Peskov, Jordan L. Boyd-Graber, Philip Resnik |
NeurIPS | 6 |
| 2020 | A Joint Model for Document Segmentation and Segment LabelingabstractText segmentation aims to uncover latent structure by dividing text from a document into coherent sections.Where previous work on text segmentation considers the tasks of document segmentation and segment labeling separately, we show that the tasks contain complementary information and are best addressed jointly.We introduce the Segment Pooling LSTM (S-LSTM) model, which is capable of jointly segmenting a document and labeling segments.In support of joint training, we develop a method for teaching the model to recover from errors by aligning the predicted and ground truth segments.We show that S-LSTM reduces segmentation error by 30% on average, while also improving segment labeling. Joe Barrow, Rajiv Jain, Vlad I. Morariu, Varun Manjunatha, Douglas W. Oard, Philip Resnik |
ACL | 6 |
| 2020 | A Prioritization Model for Suicidality Risk AssessmentabstractWe reframe suicide risk assessment from social media as a ranking problem whose goal is maximizing detection of severely at-risk individuals given the time available.Building on measures developed for resource-bounded document retrieval, we introduce a well founded evaluation paradigm, and demonstrate using an expert-annotated test collection that meaningful improvements over plausible cascade model baselines can be achieved using an approach that jointly ranks individuals and their social media posts. Han-Chin Shing, Philip Resnik, Douglas W. Oard |
ACL | 2 |
| 2020 | Improving Neural Topic Models using Knowledge DistillationabstractTopic models are often used to identify humaninterpretable topics to help make sense of large document collections.We use knowledge distillation to combine the best attributes of probabilistic topic models and pretrained transformers.Our modular method can be straightforwardly applied with any neural topic model to improve topic quality, which we demonstrate using two models having disparate architectures, obtaining state-of-the-art topic coherence.We show that our adaptable framework not only improves performance in the aggregate over all estimated topics, as is commonly reported, but also in head-to-head comparisons of aligned topics. Alexander Miserlis Hoyle, Pranav Goel 0001, Philip Resnik |
EMNLP (1) | 3 |
| 2019 | A Multilingual Topic Model for Learning Weighted Topic Links Across Corpora with Low ComparabilityabstractWeiwei Yang, Jordan Boyd-Graber, Philip Resnik. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jordan L. Boyd-Graber, Philip Resnik |
EMNLP/IJCNLP (1) | 3 |
| 2018 | Assessing Composition in Sentence Vector RepresentationsabstractAn important component of achieving language understanding is mastering the composition of sentence meaning, but an immediate challenge to solving this problem is the opacity of sentence vector representations produced by current neural sentence composition models. We present a method to address this challenge, developing tasks that directly target compositional meaning information in sentence vector representations with a high degree of precision and control. To enable the creation of these controlled tasks, we introduce a specialized sentence generation system that produces large, annotated sentence sets meeting specified syntactic, semantic and lexical constraints. We describe the details of the method and generation system, and then present results of experiments applying our method to probe for compositional information in embeddings from a number of existing sentence composition models. We find that the method is able to extract useful information about the differing capacities of these models, and we discuss the implications of our results with respect to these systems’ capturing of sentence information. We make available for public use the datasets used for these experiments, as well as the generation system. Allyson Ettinger, Ahmed Elgohary, Colin Phillips, Philip Resnik |
COLING | 4 |
| 2017 | Adapting Topic Models using Lexical Associations with Tree PriorsabstractModels work best when they are optimized taking into account the evaluation criteria that people care about.For topic models, people often care about interpretability, which can be approximated using measures of lexical association.We integrate lexical association into topic optimization using tree priors, which provide a flexible framework that can take advantage of both first order word associations and the higher-order associations captured by word embeddings.Tree priors improve topic interpretability without hurting extrinsic performance. Jordan L. Boyd-Graber, Philip Resnik |
EMNLP | 3 |
| 2016 | Learning Text Pair Similarity with Context-sensitive Autoencoders
Hadi Amiri, Philip Resnik, Jordan L. Boyd-Graber, Hal Daumé III |
ACL (1) | 2 |
| 2016 | A Discriminative Topic Model using Document Network StructureabstractDocument collections often have links between documents—citations, hyperlinks, or revisions—and which links are added is often based on topical similarity. To model these intuitions, we introduce a new topic model for documents situated within a network structure, integrating latent blocks of documents with a max-margin learning criterion for link prediction using topicand word-level features. Experiments on a scientific paper dataset and collection of webpages show that, by more robustly exploiting the rich link structure within a document network, our model improves link prediction, topic quality, and block distributions. Jordan L. Boyd-Graber, Philip Resnik |
ACL (1) | 3 |
| 2016 | Modeling N400 amplitude using vector space models of word representation
Allyson Ettinger, Naomi Feldman, Philip Resnik, Colin Phillips |
CogSci | 3 |
| 2016 | Retrofitting Sense-Specific Word Vectors Using Parallel TextabstractJauhar et al. (2015) recently proposed to learn sense-specific word representations by "retrofitting" standard distributional word representations to an existing ontology.We observe that this approach does not require an ontology, and can be generalized to any graph defining word senses and relations between them.We create such a graph using translations learned from parallel corpora.On a set of lexical semantic tasks, representations learned using parallel text perform roughly as well as those derived from WordNet, and combining the two representation types significantly improves performance. Allyson Ettinger, Philip Resnik, Marine Carpuat |
HLT-NAACL | 2 |
| 2015 | Tea Party in the House: A Hierarchical Ideal Point Topic Model and Its Application to Republican Legislators in the 112th CongressabstractViet-An Nguyen, Jordan Boyd-Graber, Philip Resnik, Kristina Miler. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Viet-An Nguyen, Jordan L. Boyd-Graber, Philip Resnik, Kristina Miler |
ACL (1) | 3 |
| 2015 | Birds of a Feather Linked Together: A Discriminative Topic Model using Link-based PriorsabstractA wide range of applications, from social media to scientific literature analysis, involve graphs in which documents are connected by links. We introduce a topic model for link prediction based on the intuition that linked documents will tend to have similar topic distributions, integrating a max-margin learning criterion and lexical term weights in the loss function. We validate our approach on the tweets from 2,000 Sina Weibo users and evaluate our model’s reconstruction of the social network. Jordan L. Boyd-Graber, Philip Resnik |
EMNLP | 3 |
| 2015 | Dialogue focus tracking for zero pronoun resolutionabstractSudha Rao, Allyson Ettinger, Hal Daumé III, Philip Resnik. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Sudha Rao, Allyson Ettinger, Hal Daumé III, Philip Resnik |
HLT-NAACL | 4 |
| 2014 | Political Ideology Detection Using Recursive Neural NetworksabstractAn individual's words often reveal their political ideology.Existing automated techniques to identify ideology from text focus on bags of words or wordlists, ignoring syntax.Taking inspiration from recent work in sentiment analysis that successfully models the compositional aspect of language, we apply a recursive neural network (RNN) framework to the task of identifying the political position evinced by a sentence.To show the importance of modeling subsentential elements, we crowdsource political annotations at a phrase and sentence level.Our model outperforms existing models on our newly annotated dataset and an existing dataset. Mohit Iyyer, Peter Enns, Jordan L. Boyd-Graber, Philip Resnik |
ACL (1) | 4 |
| 2014 | A Unified Model for Soft Linguistic Reordering Constraints in Statistical Machine TranslationabstractThis paper explores a simple and effective unified framework for incorporating soft linguistic reordering constraints into a hierarchical phrase-based translation system: 1) a syntactic reordering model that explores reorderings for context free grammar rules; and 2) a semantic reordering model that focuses on the reordering of predicate-argument structures.We develop novel features based on both models and use them as soft constraints to guide the translation process.Experiments on Chinese-English translation show that the reordering approach can significantly improve a state-of-the-art hierarchical phrase-based translation system.However, the gain achieved by the semantic reordering model is limited in the presence of the syntactic reordering model, and we therefore provide a detailed analysis of the behavior differences between the two. Yuval Marton, Philip Resnik, Hal Daumé III |
ACL (1) | 3 |
| 2014 | Sometimes Average is Best: The Importance of Averaging for Prediction using MCMC Inference in Topic ModelingabstractMarkov chain Monte Carlo (MCMC) approximates the posterior distribution of latent variable models by generating many samples and averaging over them.In practice, however, it is often more convenient to cut corners, using only a single sample or following a suboptimal averaging strategy.We systematically study different strategies for averaging MCMC samples and show empirically that averaging properly leads to significant improvements in prediction. Viet-An Nguyen, Jordan L. Boyd-Graber, Philip Resnik |
EMNLP | 3 |
| 2014 | Learning a Concept Hierarchy from Multi-labeled Documents
Viet-An Nguyen, Jordan L. Boyd-Graber, Philip Resnik, Jonathan D. Chang |
NIPS | 3 |
| 2014 | Modeling topic control to detect influence in conversations using nonparametric topic models
Viet-An Nguyen, Jordan L. Boyd-Graber, Philip Resnik, Deborah A. Cai, Jennifer E. Midberry |
Mach. Learn. | 3 |
| 2014 | Crowdsourced Monolingual TranslationabstractAn enormous potential exists for solving certain classes of computational problems through rich collaboration among crowds of humans supported by computers. Solutions to these problems used to involve human professionals, who are expensive to hire or difficult to find. Despite significant advances, fully automatic systems still have much room for improvement. Recent research has involved recruiting large crowds of skilled humans (“crowdsourcing”), but crowdsourcing solutions are still restricted by the availability of those skilled human participants. With translation, for example, professional translators incur a high cost and are not always available; machine translation systems have been greatly improved recently but still can only provide passable translation; and crowdsourced translation is limited by the availability of bilingual humans. This article describes crowdsourced monolingual translation, where monolingual translation is translation performed by monolingual people. Crowdsourced monolingual translation is a collaborative form of translation performed by two crowds of people who speak the source or the target language, respectively, with machine translation as the mediating device. This article describes a general protocol to handle crowdsourced monolingual translation and analyzes three systems that implemented the protocol. These systems were studied in various settings and were found to supply significant improvement in quality over both machine translation and monolingual editing of machine translation output (“postediting”). Philip Resnik, Benjamin B. Bederson |
ACM Trans. Comput. Hum. Interact. | 2 |
| 2013 | Online Relative Margin Maximization for Statistical Machine Translation
Vladimir Eidelman, Yuval Marton, Philip Resnik |
ACL (1) | 3 |
| 2013 | Using Topic Modeling to Improve Prediction of Neuroticism and Depression in College StudentsabstractWe investigate the value-add of topic modeling in text analysis for depression, and for neuroticism as a strongly associated personality measure.Using Pennebaker's Linguistic Inquiry and Word Count (LIWC) lexicon to provide baseline features, we show that straightforward topic modeling using Latent Dirichlet Allocation (LDA) yields interpretable, psychologically relevant "themes" that add value in prediction of clinical assessments. Philip Resnik, Anderson Garron, Rebecca Resnik |
EMNLP | 1 |
| 2013 | Modeling Syntactic and Semantic Structures in Hierarchical Phrase-based Translation
Philip Resnik, Hal Daumé III |
HLT-NAACL | 2 |
| 2013 | Argviz: Interactive Visualization of Topic Dynamics in Multi-party Conversations
Viet-An Nguyen, Yuening Hu, Jordan L. Boyd-Graber, Philip Resnik |
HLT-NAACL | 4 |
| 2013 | Lexical and Hierarchical Topic RegressionabstractInspired by a two-level theory that unifies agenda setting and ideological framing, we propose supervised hierarchical latent Dirichlet allocation (SHLDA) which jointly captures documents' multi-level topic structure and their polar response variables. Our model extends the nested Chinese restaurant process to discover a tree-structured topic hierarchy and uses both per-topic hierarchical and per-word lexical regression parameters to model the response variables. Experiments in a political domain and on sentiment analysis tasks show that SHLDA improves predictive accuracy while adding a new dimension of insight into how topics under discussion are framed. Viet-An Nguyen, Jordan L. Boyd-Graber, Philip Resnik |
NIPS | 3 |
| 2013 | Using targeted paraphrasing and monolingual crowdsourcing to improve translationabstractTargeted paraphrasing is a new approach to the problem of obtaining cost-effective, reasonable quality translation, which makes use of simple and inexpensive human computations by monolingual speakers in combination with machine translation. The key insight behind the process is that it is possible to spot likely translation errors with only monolingual knowledge of the target language, and it is possible to generate alternative ways to say the same thing (i.e., paraphrases) with only monolingual knowledge of the source language. Formal evaluation demonstrates that this approach can yield substantial improvements in translation quality, and the idea has been integrated into a broader framework for monolingual collaborative translation that produces fully accurate, fully fluent translations for a majority of sentences in a real-world translation task, with no involvement of human bilingual speakers. Philip Resnik, Olivia Buzek, Yakov Kronrod, Alexander J. Quinn, Benjamin B. Bederson |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2012 | SITS: A Hierarchical Nonparametric Model using Speaker Identity for Topic Segmentation in Multiparty Conversations
Viet-An Nguyen, Jordan L. Boyd-Graber, Philip Resnik |
ACL (1) | 3 |
| 2012 | Deploying monotrans widgets in the wildabstractIn this paper, we report our experience deploying the MonoTrans Widgets system in a public setting. Our work follows a line of crowd-sourced monolingual translation systems, and it is the first attempt to deploy such a system "in the wild". The results are promising, but we also found out that simultaneously drawing from multiple crowds with different expertise and sizes poses unique problems in the design of such crowd-sourcing systems. Philip Resnik, Yakov Kronrod, Benjamin B. Bederson |
CHI | 2 |
| 2012 | Encouraging Consistent Translation Choices
Ferhan Ture, Douglas W. Oard, Philip Resnik |
HLT-NAACL | 3 |
| 2012 | Soft syntactic constraints for Arabic-English hierarchical phrase-based translation
Yuval Marton, David Chiang 0001, Philip Resnik |
Mach. Transl. | 3 |
| 2011 | MonoTrans2: a new human computation system to support monolingual translationabstractIn this paper, we present MonoTrans2, a new user interface to support monolingual translation; that is, translation by people who speak only the source or target language, but not both. Compared to previous systems, MonoTrans2 supports multiple edits in parallel, and shorter tasks with less translation context. In an experiment translating children's books, we show that MonoTrans2 is able to substantially close the gap between machine translation and human bilingual translations. The percentage of sentences rated 5 out of 5 for fluency and adequacy by both bilingual evaluators in our study increased from 10% for Google Translate output to 68% for MonoTrans2. Benjamin B. Bederson, Philip Resnik, Yakov Kronrod |
CHI | 3 |
| 2010 | Holistic Sentiment Analysis Across Languages: Multilingual Supervised Latent Dirichlet Allocation
Jordan L. Boyd-Graber, Philip Resnik |
EMNLP | 2 |
| 2010 | Modeling Perspective Using Adaptor Grammars
Eric Hardisty, Jordan L. Boyd-Graber, Philip Resnik |
EMNLP | 3 |
| 2010 | Improving Translation via Targeted Paraphrasing
Philip Resnik, Olivia Buzek, Yakov Kronrod, Alexander J. Quinn, Benjamin B. Bederson |
EMNLP | 1 |
| 2010 | Discriminative Word Alignment with a Function Word Reordering Model
Hendra Setiawan, Chris Dyer, Philip Resnik |
EMNLP | 3 |
| 2010 | Translation by iterative collaboration between monolingual users
Benjamin B. Bederson, Philip Resnik |
Graphics Interface | 3 |
| 2010 | Context-free reordering, finite-state translation
Chris Dyer, Philip Resnik |
HLT-NAACL | 2 |
| 2010 | Generalizing Hierarchical Phrase-based Translation using Rules with Adjacent Nonterminals
Hendra Setiawan, Philip Resnik |
HLT-NAACL | 2 |
| 2010 | Exploiting syntactic relationships in a phrase-based decoder: an exploration
Tim Hunter, Philip Resnik |
Mach. Transl. | 2 |
| 2009 | Topological Ordering of Function Words in Hierarchical Phrase-based Translation
Hendra Setiawan, Min-Yen Kan, Haizhou Li 0001, Philip Resnik |
ACL/IJCNLP | 4 |
| 2009 | Improved Statistical Machine Translation Using Monolingually-Derived Paraphrases
Yuval Marton, Chris Callison-Burch, Philip Resnik |
EMNLP | 3 |
| 2009 | Estimating Semantic Distance Using Soft Semantic Constraints in Knowledge-Source - Corpus Hybrid Models
Yuval Marton, Saif M. Mohammad, Philip Resnik |
EMNLP | 3 |
| 2009 | More than Words: Syntactic Packaging and Implicit Sentiment
Stephan Greene, Philip Resnik |
HLT-NAACL | 2 |
| 2009 | Elements of a computational model for multi-party discourse: The turn-taking behavior of Supreme Court justicesabstractAbstract This work explores computational models of multi‐party discourse, using transcripts from U.S. Supreme Court oral arguments. The turn‐taking behavior of participants is treated as a supervised sequence‐labeling problem and modeled using first‐ and second‐order conditional random fields (CRFs). We specifically explore the hypothesis that discourse markers and personal references provide important features in such models. Results from a sequence prediction experiment demonstrate that incorporating these two types of features yields significant improvements in accuracy. Our experiments are couched in the broader context of developing tools to support legal scholarship, although we see other natural language processing applications as well. Timothy Hawes, Jimmy Lin, Philip Resnik |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2008 | Generalizing Word Lattice Translation
Chris Dyer, Smaranda Muresan, Philip Resnik |
ACL | 3 |
| 2008 | Soft Syntactic Constraints for Hierarchical Phrased-Based Translation
Yuval Marton, Philip Resnik |
ACL | 2 |
| 2008 | Online Large-Margin Training of Syntactic and Structural Translation Features
David Chiang 0001, Yuval Marton, Philip Resnik |
EMNLP | 3 |
| 2008 | Cross-Language Parser Adaptation between Related Languages
Daniel Zeman, Philip Resnik |
IJCNLP | 2 |
| 2007 | Evaluating a cross-cultural children's online book community: Lessons learned for sociability, usability, and cultural exchangeabstractThe use of computers for human-to-human communication among adults has been studied for many years, but using computer technology to enable children from all over the world to talk to each other has rarely been discussed by researchers. The goal of our research is to fill this gap and explore the design and evaluation of children’s cross-language online communities via a case study of the International Children’s Digital Library Communities (ICDLCommunities). This project supports the development of communities for children (ages 7–11) that form around the International Digital Children’s Library (ICDL) book collection. In this community the children can learn about each others’ cultures and make friends even if they do not speak the same language. They can also read and create stories and ask and answer questions about these. From this evaluation study we learned that: (i) children are very interested in their counterparts in other countries and a remarkable amount of communication takes place even when they do not share a common language; (ii) representing their identity online in many different forms is particularly important to children when communicating in an online community; (iii) children enjoy drawing but representing stories in a sequence of diagrams is challenging and needs support; and (iv) asking and answering questions without language is possible using graphical templates. In this paper we present our findings and make recommendations for designing children’s cross-cultural online communities. Anita Komlodi, Weimin Hou, Jennifer Preece, Allison Druin, Evan Golub, Jade Alburo, Sabrina Liao, Aaron Elkiss, Philip Resnik |
Interact. Comput. | 9 |
| 2005 | The Linguist's Search Engine: An Overview
Philip Resnik, Aaron Elkiss |
ACL | 1 |
| 2005 | Dictionary-based techniques for cross-language information retrieval
Gina-Anne Levow, Douglas W. Oard, Philip Resnik |
Inf. Process. Manag. | 3 |
| 2005 | Bootstrapping parsers via syntactic projection across parallel textsabstractBroad coverage, high quality parsers are available for only a handful of languages. A prerequisite for developing broad coverage parsers for more languages is the annotation of text with the desired linguistic representations (also known as “treebanking”). However, syntactic annotation is a labor intensive and time-consuming process, and it is difficult to find linguistically annotated text in sufficient quantities. In this article, we explore using parallel text to help solving the problem of creating syntactic annotation in more languages. The central idea is to annotate the English side of a parallel corpus, project the analysis to the second language, and then train a stochastic analyzer on the resulting noisy annotations. We discuss our background assumptions, describe an initial study on the “projectability” of syntactic relations, and then present two experiments in which stochastic parsers are developed with minimal human intervention via projection from English. Rebecca Hwa, Philip Resnik, Amy Weinberg, Clara I. Cabezas, Okan Kolak |
Nat. Lang. Eng. | 2 |
| 2004 | Inducing Frame Semantic Verb Classes from WordNet and LDOCEabstractThis paper presents SemFrame, a system that induces frame semantic verb classes from WordNet and LDOCE. Semantic frames are thought to have significant potential in resolving the paraphrase problem challenging many language-based applications.When compared to the handcrafted FrameNet, SemFrame achieves its best recall-precision balance with 83.2% recall (based on SemFrame's coverage of FrameNet frames) and 73.8% precision (based on SemFrame verbs' semantic relatedness to frame-evoking verbs). The next best performing semantic verb classes achieve 56.9% recall and 55.0% precision. Rebecca Green, Bonnie J. Dorr, Philip Resnik |
ACL | 3 |
| 2004 | Exploiting Hidden Meanings: Using Bilingual Text for Monolingual Annotation
Philip Resnik |
CICLing | 1 |
| 2003 | A Generative Probabilistic OCR Model for NLP Applications
Okan Kolak, William J. Byrne, Philip Resnik |
HLT-NAACL | 3 |
| 2003 | Desparately Seeking Cebuano
Douglas W. Oard, David S. Doermann, Bonnie J. Dorr, Daqing He, Philip Resnik, Amy Weinberg, William J. Byrne, Sanjeev Khudanpur, David Yarowsky, Anton Leuski, Philipp Koehn, Kevin Knight |
HLT-NAACL | 5 |
| 2003 | The Web as a Parallel CorpusabstractParallel corpora have become an essential resource for work in multilingual natural language processing. In this article, we report on our work using the STRAND system for mining parallel text on the World Wide Web, first reviewing the original algorithm and results and then presenting a set of significant enhancements. These enhancements include the use of supervised learning based on structural features of documents to improve classification performance, a new content-based measure of translational equivalence, and adaptation of the system to take advantage of the Internet Archive for mining parallel text from the Web on a large scale. Finally, the value of these techniques is demonstrated in the construction of a significant parallel corpus for a low-density language pair. Philip Resnik, Noah A. Smith |
Comput. Linguistics | 1 |
| 2003 | Making MIRACLEs: Interactive translingual search for Cebuano and HindiabstractSearching is inherently a user-centered process; people pose the questions for which machines seek answers, and ultimately people judge the degree to which retrieved documents meet their needs. Rapid development of interactive systems that use queries expressed in one language to search documents written in another poses five key challenges: (1) interaction design, (2) query formulation, (3) cross-language search, (4) construction of translated summaries, and (5) machine translation. This article describes the design of MIRACLE, an easily extensible system based on English queries that has previously been used to search French, German, and Spanish documents, and explains how the capabilities of MIRACLE were rapidly extended to accommodate Cebuano and Hindi. Evaluation results for the cross-language search component are presented for both languages, along with results from a brief full-system interactive experiment with Hindi. The article concludes with some observations on directions for further research on interactive cross-language information retrieval. Daqing He, Douglas W. Oard, Jianqiang Wang 0002, Dina Demner-Fushman, Kareem Darwish, Philip Resnik, Sanjeev Khudanpur, Michael Nossal, Michael Subotin, Anton Leuski |
ACM Trans. Asian Lang. Inf. Process. | 7 |
| 2002 | An Unsupervised Method for Word Sense Tagging using Parallel CorporaabstractWe present an unsupervised method for word sense disambiguation that exploits translation correspondences in parallel corpora. The technique takes advantage of the fact that cross-language lexicalizations of the same concept tend to be consistent, preserving some core element of its semantics, and yet also variable, reflecting differing translator preferences and the influence of context. Working with parallel corpora introduces an extra complication for evaluation, since it is difficult to find a corpus that is both sense tagged and parallel with another language; therefore we use pseudo-translations, created by machine translation systems, in order to make possible the evaluation of the approach against a standard test set. The results demonstrate that word-level translation correspondences are a valuable source of information for sense disambiguation. Mona T. Diab, Philip Resnik |
ACL | 2 |
| 2002 | Evaluating Translational Correspondence using Annotation ProjectionabstractRecently, statistical machine translation models have begun to take advantage of higher level linguistic structures such as syntactic dependencies. Underlying these models is an assumption about the directness of translational correspondence between sentences in the two languages; however, the extent to which this assumption is valid and useful is not well understood. In this paper, we present an empirical study that quantifies the degree to which syntactic dependencies are preserved when parses are projected directly from English to Chinese. Our results show that although the direct correspondence assumption is often too restrictive, a small set of principled, elementary linguistic transformations can boost the quality of the projected Chinese parses by 76% relative to the unimproved baseline. Rebecca Hwa, Philip Resnik, Amy Weinberg, Okan Kolak |
ACL | 2 |
| 2001 | Mapping Lexical Entries in a Verbs Database to WordNet SensesabstractThis paper describes automatic techniques for mapping 9611 entries in a database of English verbs to WordNet senses. The verbs were initially grouped into 491 classes based on syntactic features. Mapping these verbs into WordNet senses provides a resource that supports disambiguation in multilingual applications such as machine translation and cross-language information retrieval. Our techniques make use of (1) a training set of 1791 disambiguated entries, representing 1442 verb entries from 167 classes; (2) word sense probabilities, from frequency counts in a tagged corpus; (3) semantic similarity of WordNet senses for verbs within the same class; (4) probabilistic correlations between WordNet data and attributes of the verb classes. The best results achieved 72% precision and 58% recall, versus a lower bound of 62% precision and 38% recall for assigning the most frequently occurring WordNet sense, and an upper bound of 87% precision and 75% recall for human judgment. Rebecca Green, Lisa Pearl, Bonnie J. Dorr, Philip Resnik |
ACL | 4 |
| 1999 | Mining the Web for Bilingual TextabstractSTRAND (Resnik, 1998) is a language-independent system for automatic discovery of text in parallel translation on the World Wide Web. This paper extends the preliminary STRAND results by adding automatic language identification, scaling up by orders of magnitude, and formally evaluating performance. The most recent end-product is an automatically acquired parallel corpus comprising 2491 English-French document pairs, approximately 1.5 million words per language. Philip Resnik |
ACL | 1 |
| 1999 | Support for Interactive Document Selection in Cross-Language Information Retrieval
Douglas W. Oard, Philip Resnik |
Inf. Process. Manag. | 2 |
| 1999 | Semantic Similarity in a Taxonomy: An Information-Based Measure and its Application to Problems of Ambiguity in Natural LanguageabstractThis article presents a measure of semantic similarity in an IS-A taxonomy based on the notion of shared information content. Experimental evaluation against a benchmark set of human similarity judgments demonstrates that the measure performs better than the traditional edge-counting approach. The article presents algorithms that take advantage of taxonomic similarity in resolving syntactic and semantic ambiguity, along with experimental results demonstrating their effectiveness. Philip Resnik |
J. Artif. Intell. Res. | 1 |
| 1999 | Distinguishing systems and distinguishing senses: new evaluation methods for Word Sense DisambiguationabstractResnik and Yarowsky (1997) made a set of observations about the state-of-the-art in automatic word sense disambiguation and, motivated by those observations, offered several specific proposals regarding improved evaluation criteria, common training and testing resources, and the definition of sense inventories. Subsequent discussion of those proposals resulted in SENSEVAL, the first evaluation exercise for word sense disambiguation (Kilgarriff and Palmer 2000). This article is a revised and extended version of our 1997 workshop paper, reviewing its observations and proposals and discussing them in light of the SENSEVAL exercise. It also includes a new in-depth empirical study of translingually-based sense inventories and distance measures, using statistics collected from native-speaker annotations of 222 polysemous contexts across 12 languages. These data show that monolingual sense distinctions at most levels of granularity can be effectively captured by translations into some set of second languages, especially as language family distance increases. In addition, the probability that a given sense pair will tend to lexicalize differently across languages is shown to correlate with semantic salience and sense granularity; sense hierarchies automatically generated from such distance matrices yield results remarkably similar to those created by professional monolingual lexicographers. Philip Resnik, David Yarowsky |
Nat. Lang. Eng. | 1 |
| 1995 | Using Information Content to Evaluate Semantic Similarity in a Taxonomy
Philip Resnik |
IJCAI | 1 |
| 1994 | A Rule-Based Approach to Prepositional Phrase Attachment Disambiguation
Eric Brill, Philip Resnik |
COLING | 2 |
| 1992 | A Class-Based Approach to Lexical Discoveryabstractthis paper I propose a generalization of lexical association techniques that is intended to facilitate statistical discovery of facts involving word classes rather than individual words. Although defining association measures over classes (as sets of words) is straightforward in theory, making direct use of such a definition is impractical because there are simply too many classes to consider. Rather than considering all possible classes, I propose constraining the set of possible word classes by using a broad-coverage lexical/conceptual hierarchy [Miller, 1990] Philip Resnik |
ACL | 1 |
| 1992 | Left-Corner Parsing And Psychological Plausibility
Philip Resnik |
COLING | 1 |
| 1992 | Probabilistic Tree-Adjoining Grammar As A Framework For Statistical Natural Language Processing
Philip Resnik |
COLING | 1 |
| 1990 | Multiple Underlying Systems: Translating User Requests into Programs to Produce AnswersabstractA user may typically need to combine the strengths of more than one system in order to perform a task. In this paper, we describe a component of the Janus natural language interface that translates intensional logic expressions representing the meaning of a request into executable code for each application program, chooses which combination of application systems to use, and designs the transfer of data among them in order to provide an answer. The complete Janus natural language system has been ported to two large command and control decision support aids. Robert J. Bobrow, Philip Resnik, Ralph M. Weischedel |
ACL | 2 |