Jaime G. Carbonell

dblp:56/3395 · DBLP profile ↗
← Back
215ranked-venue papers
32as first author
2since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 159 · 29 first-author · 2 since 2021Databases, data management, data science and information retrieval · 49 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 35 · 8 first-authorApplied, interdisciplinary, general and emerging computing · 28Human-computer interaction and ubiquitous computing · 18Theory of computation · 4 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
63 papers
Information extraction and text analysis · 36% Representation and self-supervised learning · 14% Transfer learning and domain adaptation · 11%
Databases, data mining, and information retrieval
25 papers
Information retrieval · 58% Data mining · 17% Machine learning and data management · 9%
Interdisciplinary, comprehensive, and emerging computing
12 papers
Bioinformatics and computational biology · 78% Computational science and engineering · 11% Computing education · 9%
Theoretical computer science
4 papers
Mathematical optimization · 68% Algorithms and data structures · 32% Automated reasoning and model checking · 0%

Topics — the 30 heaviest of 183, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
named entity recognition
1.762020
Soft Gazetteers for Low-Resource Named Entity Recognition · ACL 2020
A Little Annotation does a Lot of Good: A Study in Bootstrapping Low-resource Named Entity Recognizers · EMNLP/IJCNLP (1) 2019
Neural Cross-lingual Named Entity Recognition with Minimal Resources · EMNLP 2018
Natural language and speech › Information extraction and text analysis › named entity recognition
low-resource named entity recognition
1.442020
Soft Gazetteers for Low-Resource Named Entity Recognition · ACL 2020
A Little Annotation does a Lot of Good: A Study in Bootstrapping Low-resource Named Entity Recognizers · EMNLP/IJCNLP (1) 2019
Adapting Word Embeddings to New Languages with Morphological and Phonological Subword Representations · EMNLP 2018
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
1.132020
Cross-lingual Alignment vs Joint Training: A Comparative Study and A Simple Unified Framework · ICLR 2020
A Little Annotation does a Lot of Good: A Study in Bootstrapping Low-resource Named Entity Recognizers · EMNLP/IJCNLP (1) 2019
Phonologically Aware Neural Model for Named Entity Recognition in Low Resource Transfer Settings · EMNLP 2016
Machine learning › Learning paradigms
multi-task learning
0.942017
Active Learning from Peers · NIPS 2017
Self-Paced Multitask Learning with Shared Knowledge · IJCAI 2017
Adaptive Smoothed Online Multi-Task Learning · NIPS 2016
Natural language and speech › Information extraction and text analysis
semantic role labeling
0.832019
Towards Semi-Supervised Learning for Deep Semantic Role Labeling · EMNLP 2018
DeepCx: A transition-based approach for shallow semantic parsing with complex constructional triggers · EMNLP 2018
Gradient-Based Inference for Networks with Output Constraints · AAAI 2019
Bioinformatics and computational biology
protein-protein interaction prediction
0.742016
Multitask Matrix Completion for Learning Protein Interactions Across Diseases · RECOMB 2016
Multitask learning for host-pathogen protein interactions · Bioinform. 2013
Techniques to cope with missing data in host-pathogen protein interaction prediction · Bioinform. 2012
Machine learning › Learning paradigms
semi-supervised learning
0.632020
Semi-Supervised Learning on Meta Structure: Multi-Task Tagging and Parsing in Low-Resource Scenarios · AAAI 2020
Towards Semi-Supervised Learning for Deep Semantic Role Labeling · EMNLP 2018
Graph-Based Semi-Supervised Learning as a Generative Model · IJCAI 2007
Machine learning › Learning paradigms › multi-task learning
online multi-task learning
0.522017
Active Learning from Peers · NIPS 2017
Adaptive Smoothed Online Multi-Task Learning · NIPS 2016
Natural language and speech › Information extraction and text analysis
semantic parsing
0.522018
DeepCx: A transition-based approach for shallow semantic parsing with complex constructional triggers · EMNLP 2018
A Discriminative Graph-Based Parser for the Abstract Meaning Representation · ACL (1) 2014
Natural language and speech › Information extraction and text analysis › entity linking
cross-lingual entity linking
0.522020
Zero-Shot Neural Transfer for Cross-Lingual Entity Linking · AAAI 2019
Soft Gazetteers for Low-Resource Named Entity Recognition · ACL 2020
Machine learning › Representation and self-supervised learning
domain-specific representation
0.512021
Domain Adaptation with Invariant Representation Learning: What Transformations to Learn? · NeurIPS 2021
Machine learning › Representation and self-supervised learning › representation learning
invariant representation learning
0.512021
Domain Adaptation with Invariant Representation Learning: What Transformations to Learn? · NeurIPS 2021
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.512021
Domain Adaptation with Invariant Representation Learning: What Transformations to Learn? · NeurIPS 2021
Machine learning › Efficient and distributed learning
active learning
0.522017
Active Learning from Peers · NIPS 2017
Buy-in-Bulk Active Learning · NIPS 2013
Natural language and speech › Information extraction and text analysis › text classification
active learning for text classification
0.412020
Voice for the Voiceless: Active Sampling to Detect Comments Supporting the Rohingyas · AAAI 2020
Natural language and speech › Information extraction and text analysis › multilingual NLP
code-switching
0.412020
Harnessing Code Switching to Transcend the Linguistic Barrier · IJCAI 2020
Machine learning › Graph learning › limited supervision › multi-view semi-supervised learning
co-training
0.412020
Semi-Supervised Learning on Meta Structure: Multi-Task Tagging and Parsing in Low-Resource Scenarios · AAAI 2020
Machine learning › Representation and self-supervised learning › representation matching › feature alignment › embedding alignment
cross-lingual alignment
0.412020
Cross-lingual Alignment vs Joint Training: A Comparative Study and A Simple Unified Framework · ICLR 2020
Machine learning › Representation and self-supervised learning › text embedding › text representation learning
cross-lingual representation learning
0.412020
Cross-lingual Alignment vs Joint Training: A Comparative Study and A Simple Unified Framework · ICLR 2020
Machine learning › Efficient and distributed learning
data selection
0.412020
Optimizing Data Usage via Differentiable Rewards · ICML 2020
Natural language and speech › Information extraction and text analysis › syntactic parsing
dependency parsing
0.412020
Semi-Supervised Learning on Meta Structure: Multi-Task Tagging and Parsing in Low-Resource Scenarios · AAAI 2020
Natural language and speech › Information extraction and text analysis › abusive language detection
hate speech detection
0.412020
Voice for the Voiceless: Active Sampling to Detect Comments Supporting the Rohingyas · AAAI 2020
Machine learning › Learning paradigms
lifelong learning
0.412020
Efficient Meta Lifelong-Learning with Limited Memory · EMNLP (1) 2020
Natural language and speech › Information extraction and text analysis › syntactic parsing › dependency parsing
low-resource dependency parsing
0.412020
Semi-Supervised Learning on Meta Structure: Multi-Task Tagging and Parsing in Low-Resource Scenarios · AAAI 2020
Natural language and speech › Information extraction and text analysis › sequence labeling › part-of-speech tagging
low-resource part-of-speech tagging
0.412020
Semi-Supervised Learning on Meta Structure: Multi-Task Tagging and Parsing in Low-Resource Scenarios · AAAI 2020
Natural language and speech › Language models and text generation
multilingual language models
0.412020
Harnessing Code Switching to Transcend the Linguistic Barrier · IJCAI 2020
Machine learning › Representation and self-supervised learning
multi-view learning
0.412020
Semi-Supervised Learning on Meta Structure: Multi-Task Tagging and Parsing in Low-Resource Scenarios · AAAI 2020
Natural language and speech › Information extraction and text analysis › sequence labeling
part-of-speech tagging
0.412020
Semi-Supervised Learning on Meta Structure: Multi-Task Tagging and Parsing in Low-Resource Scenarios · AAAI 2020
Machine learning › Reinforcement learning
reward learning
0.412020
Optimizing Data Usage via Differentiable Rewards · ICML 2020
Natural language and speech › Information extraction and text analysis
text classification
0.412020
Voice for the Voiceless: Active Sampling to Detect Comments Supporting the Rohingyas · AAAI 2020

Methods — techniques the papers use, named apart from their topics

active learning · 1.2neural network · 1.1comment embeddings · 0.9attention · 0.5multi-task learning · 0.5causal mechanism invariance · 0.5graph inference · 0.4nearest-neighbor sampling · 0.4nearest neighbor sampling · 0.4multi-view learning · 0.4cross-lingual entity linking · 0.4consensus promotion · 0.4co-training · 0.4re-ranking · 0.3multimodal fusion · 0.3matrix completion · 0.2mirror prox · 0.2accelerated primal-dual algorithm · 0.2
YearPublicationVenuePosition
2021 StructSum: Summarization via Structured Representations
abstract
Vidhisha Balachandran, Artidoro Pagnoni, Jay Yoon Lee, Dheeraj Rajagopal, Jaime Carbonell, Yulia Tsvetkov. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Vidhisha Balachandran, Artidoro Pagnoni, Jay-Yoon Lee, Dheeraj Rajagopal, Jaime G. Carbonell, Yulia Tsvetkov
EACL5
2021 Domain Adaptation with Invariant Representation Learning: What Transformations to Learn?
abstract
Unsupervised domain adaptation, as a prevalent transfer learning setting, spans many real-world applications. With the increasing representational power and applicability of neural networks, state-of-the-art domain adaptation methods make use of deep architectures to map the input features $X$ to a latent representation $Z$ that has the same marginal distribution across domains. This has been shown to be insufficient for generating optimal representation for classification, and to find conditionally invariant representations, usually strong assumptions are needed. We provide reasoning why when the supports of the source and target data from overlap, any map of $X$ that is fixed across domains may not be suitable for domain adaptation via invariant features. Furthermore, we develop an efficient technique in which the optimal map from $X$ to $Z$ also takes domain-specific information as input, in addition to the features $X$. By using the property of minimal changes of causal mechanisms across domains, our model also takes into account the domain-specific information to ensure that the latent representation $Z$ does not discard valuable information about $Y$. We demonstrate the efficacy of our method via synthetic and real-world data experiments. The code is available at: \texttt{https://github.com/DMIRLAB-Group/DSAN}.
Petar Stojanov, Zijian Li 0001, Mingming Gong, Ruichu Cai, Jaime G. Carbonell, Kun Zhang 0001
NeurIPS5
2020 Semi-Supervised Learning on Meta Structure: Multi-Task Tagging and Parsing in Low-Resource Scenarios
abstract
Multi-view learning makes use of diverse models arising from multiple sources of input or different feature subsets for the same task. For example, a given natural language processing task can combine evidence from models arising from character, morpheme, lexical, or phrasal views. The most common strategy with multi-view learning, especially popular in the neural network community, is to unify multiple representations into one unified vector through concatenation, averaging, or pooling, and then build a single-view model on top of the unified representation. As an alternative, we examine whether building one model per view and then unifying the different models can lead to improvements, especially in low-resource scenarios. More specifically, taking inspiration from co-training methods, we propose a semi-supervised learning approach based on multi-view models through consensus promotion, and investigate whether this improves overall performance. To test the multi-view hypothesis, we use moderately low-resource scenarios for nine languages and test the performance of the joint model for part-of-speech tagging and dependency parsing. The proposed model shows significant improvements across the test cases, with average gains of -0.9 ∼ +9.3 labeled attachment score (LAS) points. We also investigate the effect of unlabeled data on the proposed model by varying the amount of training data and by using different domains of unlabeled data.
Jay-Yoon Lee, Jaime G. Carbonell, Thierry Poibeau
AAAI3
2020 Voice for the Voiceless: Active Sampling to Detect Comments Supporting the Rohingyas
abstract
The Rohingya refugee crisis is one of the biggest humanitarian crises of modern times with more than 700,000 Rohingyas rendered homeless according to the United Nations High Commissioner for Refugees. While it has received sustained press attention globally, no comprehensive research has been performed on social media pertaining to this large evolving crisis. In this work, we construct a substantial corpus of YouTube video comments (263,482 comments from 113,250 users in 5,153 relevant videos) with an aim to analyze the possible role of AI in helping a marginalized community. Using a novel combination of multiple Active Learning strategies and a novel active sampling strategy based on nearest-neighbors in the comment-embedding space, we construct a classifier that can detect comments defending the Rohingyas among larger numbers of disparaging and neutral ones. We advocate that beyond the burgeoning field of hate speech detection, automatic detection of help speech can lend voice to the voiceless people and make the internet safer for marginalized communities.
Shriphani Palakodety, Ashiqur R. KhudaBukhsh, Jaime G. Carbonell
AAAI3
2020 Soft Gazetteers for Low-Resource Named Entity Recognition
abstract
Traditional named entity recognition models use gazetteers (lists of entities) as features to improve performance.Although modern neural network models do not require such handcrafted features for strong performance, recent work (Wu et al., 2018) has demonstrated their utility for named entity recognition on English data.However, designing such features for low-resource languages is challenging, because exhaustive entity gazetteers do not exist in these languages.To address this problem, we propose a method of "soft gazetteers" that incorporates ubiquitously available information from English knowledge bases, such as Wikipedia, into neural named entity recognition models through cross-lingual entity linking.Our experiments on four low-resource languages show an average improvement of 4 points in F1 score. 1
Shruti Rijhwani, Shuyan Zhou, Graham Neubig, Jaime G. Carbonell
ACL4
2020 Hope Speech Detection: A Computational Analysis of the Voice of Peace
abstract
The recent Pulwama terror attack (February 14, 2019, Pulwama, Kashmir) triggered a chain of escalating events between India and Pakistan adding another episode to their 70-year-old dispute over Kashmir. The present era of ubiquitious social media has never seen nuclear powers closer to war. In this paper, we analyze this evolving international crisis via a substantial corpus constructed using comments on YouTube videos (921,235 English comments posted by 392,460 users out of 2.04 million overall comments by 791,289 users on 2,890 videos). Our main contributions in the paper are three-fold. First, we present an observation that polyglot word-embeddings reveal precise and accurate language clusters, and subsequently construct a document language-identification technique with negligible annotation requirements. We demonstrate the viability and utility across a variety of data sets involving several low-resource languages. Second, we present an analysis on temporal trends of pro-peace and pro-war intent observing that when tensions between the two nations were at their peak, pro-peace intent in the corpus was at its highest point. Finally, in the context of heated discussions in a politically tense situation where two nations are at the brink of a full-fledged war, we argue the importance of automatic identification of user-generated web content that can diffuse hostility and address this prediction task, dubbed \emph{hope-speech detection}.
Shriphani Palakodety, Ashiqur R. KhudaBukhsh, Jaime G. Carbonell
ECAI3
2020 Mining Insights from Large-Scale Corpora Using Fine-Tuned Language Models
abstract
Mining insights from large volume of social media texts with minimal supervision is a highly challenging Natural Language Processing (NLP) task. While Language Models' (LMs) efficacy in several downstream tasks is well-studied, assessing their applicability in answering relational questions, tracking perception or mining deeper insights is under-explored. Few recent lines of work have scratched the surface by studying pre-trained LMs' (e.g., BERT) capability in answering relational questions through "fill-in-the-blank" cloze statements (e.g., [Dante was born in MASK]). BERT predicts the MASK-ed word with a list of words ranked by probability (in this case, BERT successfully predicts Florence with the highest probability). In this paper, we conduct a feasibility study of fine-tuned LMs with a different focus on tracking polls, tracking community perception and mining deeper insights typically obtained through costly surveys. Our main focus is on a substantial corpus of video comments extracted from YouTube videos (6,182,868 comments on 130,067 videos by 1,518,077 users) posted within 100 days prior to the 2019 Indian General Election. Using fill-in-the-blank cloze statements against a recent high-performance language modeling algorithm, BERT, we present a novel application of this family of tools that is able to (1) aggregate political sentiment (2) reveal community perception and (3) track evolving national priorities and issues of interest.
Shriphani Palakodety, Ashiqur R. KhudaBukhsh, Jaime G. Carbonell
ECAI3
2020 The Refugee Experience Online: Surfacing Positivity Amidst Hate
abstract
How can Artificial Intelligence help a stateless minority from online abuse? Research efforts in hate speech detection thus far have largely focused on identifying and subsequently filtering out negative content that specifically targets them. In this paper, we highlight a recent work [8] which tackles a different aspect of web-vulnerability of marginalized communities: sparsity of prominority voices championing their cause. The highlighted paper advocates that blocking hate alone may not be sufficient in these cases as the internet shapes community perception to a great extent in modern times and supportive comments to a vulnerable community serve a different purpose. Using an Active Sampling approach, the paper constructs a nuanced voice-for-the-voiceless classifier that automatically discovers comments supporting a (allegedly) persecuted minority. In the context of the Rohingya refugee crisis, one of the biggest humanitarian crises of modern times, the paper presents promising results that can substantially aid content moderation efforts in finding positive content supporting the Rohingyas.
Shriphani Palakodety, Ashiqur R. KhudaBukhsh, Jaime G. Carbonell
ECAI3
2020 Minimizing and Recovering from the Effect of Concept Drift via Feature Selection
Daegun Won, Peter J. Jansen, Jaime G. Carbonell
ECAI3
2020 Efficient Meta Lifelong-Learning with Limited Memory
abstract
Current natural language processing models work well on a single task, yet they often fail to continuously learn new tasks without forgetting previous ones as they are re-trained throughout their lifetime, a challenge known as lifelong learning.State-of-the-art lifelong language learning methods store past examples in episodic memory and replay them at both training and inference time.However, as we show later in our experiments, there are three significant impediments: (1) needing unrealistically large memory module to achieve good performance, (2) suffering from negative transfer, (3) requiring multiple local adaptation steps for each test example that significantly slows down the inference speed.In this paper, we identify three common principles of lifelong learning methods and propose an efficient meta-lifelong framework that combines them in a synergistic fashion.To achieve sample efficiency, our method trains the model in a manner that it learns a better initialization for local adaptation.Extensive experiments on text classification and question answering benchmarks demonstrate the effectiveness of our framework by achieving state-of-the-art performance using merely 1% memory size and narrowing the gap with multi-task learning.We further show that our method alleviates both catastrophic forgetting and negative transfer at the same time.
Sanket Vaibhav Mehta, Barnabás Póczos, Jaime G. Carbonell
EMNLP (1)4
2020 Cross-lingual Alignment vs Joint Training: A Comparative Study and A Simple Unified Framework
Jiateng Xie, Ruochen Xu, Yiming Yang 0002, Graham Neubig, Jaime G. Carbonell
ICLR6
2020 Optimizing Data Usage via Differentiable Rewards
abstract
To acquire a new skill, humans learn better and faster if a tutor, based on their current knowledge level, informs them of how much attention they should pay to particular content or practice problems. Similarly, a machine learning model could potentially be trained better with a scorer that “adapts” to its current learning state and estimates the importance of each training data instance. Training such an adaptive scorer efficiently is a challenging problem; in order to precisely quantify the effect of a data instance at a given time during the training, it is typically necessary to first complete the entire training process. To efficiently optimize data usage, we propose a reinforcement learning approach called Differentiable Data Selection (DDS). In DDS, we formulate a scorer network as a learnable function of the training data, which can be efficiently updated along with the main model being trained. Specifically, DDS updates the scorer with an intuitive reward signal: it should up-weigh the data that has a similar gradient with a dev set upon which we would finally like to perform well. Without significant computing overhead, DDS delivers strong and consistent improvements over several strong baselines on two very different tasks of machine translation and image classification.
Xinyi Wang 0001, Paul Michel, Antonios Anastasopoulos, Jaime G. Carbonell, Graham Neubig
ICML5
2020 Harnessing Code Switching to Transcend the Linguistic Barrier
abstract
Code mixing (or code switching) is a common phenomenon observed in social-media content generated by a linguistically diverse user-base. Studies show that in the Indian sub-continent, a substantial fraction of social media posts exhibit code switching. While the difficulties posed by code mixed documents to further downstream analyses are well-understood, lending visibility to code mixed documents under certain scenarios may have utility that has been previously overlooked. For instance, a document written in a mixture of multiple languages can be partially accessible to a wider audience; this could be particularly useful if a considerable fraction of the audience lacks fluency in one of the component languages. In this paper, we provide a systematic approach to sample code mixed documents leveraging a polyglot embedding based method that requires minimal supervision. In the context of the 2019 India-Pakistan conflict triggered by the Pulwama terror attack, we demonstrate an untapped potential of harnessing code mixing for human well-being: starting from an existing hostility diffusing hope speech classifier solely trained on English documents, code mixed documents are utilized to perform cross-lingual sampling and retrieve hope speech content written in a low-resource but widely used language - Romanized Hindi. Our proposed pipeline requires minimal supervision and holds promise in substantially reducing web moderation efforts. A further exploratory study on a new COVID-19 data set introduced in this paper demonstrates the generalizability of our cross-lingual sampling technique.
Ashiqur R. KhudaBukhsh, Shriphani Palakodety, Jaime G. Carbonell
IJCAI3
2020 Improving Candidate Generation for Low-resource Cross-lingual Entity Linking
abstract
Cross-lingual entity linking (XEL) is the task of finding referents in a target-language knowledge base (KB) for mentions extracted from source-language texts. The first step of (X)EL is candidate generation, which retrieves a list of plausible candidate entities from the target-language KB for each mention. Approaches based on resources from Wikipedia have proven successful in the realm of relatively high-resource languages, but these do not extend well to low-resource languages with few, if any, Wikipedia pages. Recently, transfer learning methods have been shown to reduce the demand for resources in the low-resource languages by utilizing resources in closely related languages, but the performance still lags far behind their high-resource counterparts. In this paper, we first assess the problems faced by current entity candidate generation methods for low-resource XEL, then propose three improvements that (1) reduce the disconnect between entity mentions and KB entries, and (2) improve the robustness of the model to low-resource scenarios. The methods are simple, but effective: We experiment with our approach on seven XEL datasets and find that they yield an average gain of 16.9% in Top-30 gold candidate recall, compared with state-of-the-art baselines. Our improved model also yields an average gain of 7.9% in in-KB accuracy of end-to-end XEL. 1
Shuyan Zhou, Shruti Rijhwani, John Wieting, Jaime G. Carbonell, Graham Neubig
Trans. Assoc. Comput. Linguistics4
2019 Gradient-Based Inference for Networks with Output Constraints
abstract
Practitioners apply neural networks to increasingly complex problems in natural language processing, such as syntactic parsing and semantic role labeling that have rich output structures. Many such structured-prediction problems require deterministic constraints on the output values; for example, in sequence-to-sequence syntactic parsing, we require that the sequential outputs encode valid trees. While hidden units might capture such properties, the network is not always able to learn such constraints from the training data alone, and practitioners must then resort to post-processing. In this paper, we present an inference method for neural networks that enforces deterministic constraints on outputs without performing rule-based post-processing or expensive discrete search. Instead, in the spirit of gradient-based training, we enforce constraints with gradient-based inference (GBI): for each input at test-time, we nudge continuous model weights until the network’s unconstrained inference procedure generates an output that satisfies the constraints. We study the efficacy of GBI on three tasks with hard constraints: semantic role labeling, syntactic parsing, and sequence transduction. In each case, the algorithm not only satisfies constraints, but improves accuracy, even when the underlying network is stateof-the-art.
Jay-Yoon Lee, Sanket Vaibhav Mehta, Michael L. Wick, Jean-Baptiste Tristan, Jaime G. Carbonell
AAAI5
2019 Zero-Shot Neural Transfer for Cross-Lingual Entity Linking
abstract
Cross-lingual entity linking maps an entity mention in a source language to its corresponding entry in a structured knowledge base that is in a different (target) language. While previous work relies heavily on bilingual lexical resources to bridge the gap between the source and the target languages, these resources are scarce or unavailable for many low-resource languages. To address this problem, we investigate zero-shot cross-lingual entity linking, in which we assume no bilingual lexical resources are available in the source low-resource language. Specifically, we propose pivot-basedentity linking, which leverages information from a highresource “pivot” language to train character-level neural entity linking models that are transferred to the source lowresource language in a zero-shot manner. With experiments on 9 low-resource languages and transfer through a total of54 languages, we show that our proposed pivot-based framework improves entity linking accuracy 17% (absolute) on average over the baseline systems, for the zero-shot scenario.1 Further, we also investigate the use of language-universal phonological representations which improves average accuracy (absolute) by 36% when transferring between languages that use different scripts.
Shruti Rijhwani, Jiateng Xie, Graham Neubig, Jaime G. Carbonell
AAAI4
2019 Transformer-XL: Attentive Language Models beyond a Fixed-Length Context
abstract
Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling.We propose a novel neural architecture Transformer-XL that enables learning dependency beyond a fixed length without disrupting temporal coherence.It consists of a segment-level recurrence mechanism and a novel positional encoding scheme.Our method not only enables capturing longer-term dependency, but also resolves the context fragmentation problem.As a result, Transformer-XL learns dependency that is 80% longer than RNNs and 450% longer than vanilla Transformers, achieves better performance on both short and long sequences, and is up to 1,800+ times faster than vanilla Transformers during evaluation.Notably, we improve the state-ofthe-art results of bpc/perplexity to 0.99 on en-wiki8, 1.08 on text8, 18.3 on WikiText-103, 21.8 on One Billion Word, and 54.5 on Penn Treebank (without finetuning).When trained only on WikiText-103, Transformer-XL manages to generate reasonably coherent, novel text articles with thousands of tokens.Our code, pretrained models, and hyperparameters are available in both Tensorflow and PyTorch 1 .
Zihang Dai, Zhilin Yang 0001, Yiming Yang 0002, Jaime G. Carbonell, Quoc V. Le, Ruslan Salakhutdinov
ACL (1)4
2019 Domain Adaptation of Neural Machine Translation by Lexicon Induction
abstract
It has been previously noted that neural machine translation (NMT) is very sensitive to domain shift.In this paper, we argue that this is a dual effect of the highly lexicalized nature of NMT, resulting in failure for sentences with large numbers of unknown words, and lack of supervision for domain-specific words.To remedy this problem, we propose an unsupervised adaptation method which finetunes a pre-trained out-of-domain NMT model using a pseudo-in-domain corpus.Specifically, we perform lexicon induction to extract an in-domain lexicon, and construct a pseudo-parallel in-domain corpus by performing word-for-word back-translation of monolingual in-domain target sentences.In five domains over twenty pairwise adaptation settings and two model architectures, our method achieves consistent improvements without using any in-domain parallel sentences, improving up to 14 BLEU over unadapted models, and up to 2 BLEU over strong back-translation baselines.
Junjie Hu 0001, Mengzhou Xia, Graham Neubig, Jaime G. Carbonell
ACL (1)4
2019 Low-Dimensional Density Ratio Estimation for Covariate Shift Correction
abstract
Covariate shift is a prevalent setting for supervised learning in the wild when the training and test data are drawn from different time periods, different but related domains, or via different sampling strategies. This paper addresses a transfer learning setting, with covariate shift between source and target domains. Most existing methods for correcting covariate shift exploit density ratios of the features to reweight the source-domain data, and when the features are high-dimensional, the estimated density ratios may suffer large estimation variances, leading to poor performance of prediction under covariate shift. In this work, we investigate the dependence of covariate shift correction performance on the dimensionality of the features, and propose a correction method that finds a low-dimensional representation of the features, which takes into account feature relevant to the target $Y$, and exploits the density ratio of this representation for importance reweighting. We discuss the factors that affect the performance of our method, and demonstrate its capabilities on both pseudo-real data and real-world applications.
Petar Stojanov, Mingming Gong, Jaime G. Carbonell, Kun Zhang 0001
AISTATS3
2019 Data-Driven Approach to Multiple-Source Domain Adaptation
abstract
A key problem in domain adaptation is determining what to transfer across different domains. We propose a data-driven method to represent these changes across multiple source domains and perform unsupervised domain adaptation. We assume that the joint distributions follow a specific generating process and have a small number of identifiable changing parameters, and develop a data-driven method to identify the changing parameters by learning low-dimensional representations of the changing class-conditional distributions across multiple source domains. The learned low-dimensional representations enable us to reconstruct the target-domain joint distribution from unlabeled target-domain data, and further enable predicting the labels in the target domain. We demonstrate the efficacy of this method by conducting experiments on synthetic and real datasets.
Petar Stojanov, Mingming Gong, Jaime G. Carbonell, Kun Zhang 0001
AISTATS3
2019 Characterizing and Avoiding Negative Transfer
abstract
When labeled data is scarce for a specific target task, transfer learning often offers an effective solution by utilizing data from a related source task. However, when transferring knowledge from a less related source, it may inversely hurt the target performance, a phenomenon known as negative transfer. Despite its pervasiveness, negative transfer is usually described in an informal manner, lacking rigorous definition, careful analysis, or systematic treatment. This paper proposes a formal definition of negative transfer and analyzes three important aspects thereof. Stemming from this analysis, a novel technique is proposed to circumvent negative transfer by filtering out unrelated source data. Based on adversarial networks, the technique is highly generic and can be applied to a wide range of transfer learning algorithms. The proposed approach is evaluated on six state-of-the-art deep transfer methods via experiments on four benchmark datasets with varying levels of difficulty. Empirically, the proposed method consistently improves the performance of all baseline methods and largely avoids negative transfer, even when the source data is degenerate.
Zihang Dai, Barnabás Póczos, Jaime G. Carbonell
CVPR4
2019 A Little Annotation does a Lot of Good: A Study in Bootstrapping Low-resource Named Entity Recognizers
abstract
Aditi Chaudhary, Jiateng Xie, Zaid Sheikh, Graham Neubig, Jaime Carbonell. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Aditi Chaudhary, Jiateng Xie, Zaid Sheikh, Graham Neubig, Jaime G. Carbonell
EMNLP/IJCNLP (1)5
2019 Learning Rhyming Constraints using Structured Adversaries
abstract
Harsh Jhamtani, Sanket Vaibhav Mehta, Jaime Carbonell, Taylor Berg-Kirkpatrick. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Harsh Jhamtani, Sanket Vaibhav Mehta, Jaime G. Carbonell, Taylor Berg-Kirkpatrick
EMNLP/IJCNLP (1)3
2019 XLNet: Generalized Autoregressive Pretraining for Language Understanding
abstract
With the capability of modeling bidirectional contexts, denoising autoencoding based pretraining like BERT achieves better performance than pretraining approaches based on autoregressive language modeling. However, relying on corrupting the input with masks, BERT neglects dependency between the masked positions and suffers from a pretrain-finetune discrepancy. In light of these pros and cons, we propose XLNet, a generalized autoregressive pretraining method that (1) enables learning bidirectional contexts by maximizing the expected likelihood over all permutations of the factorization order and (2) overcomes the limitations of BERT thanks to its autoregressive formulation. Furthermore, XLNet integrates ideas from Transformer-XL, the state-of-the-art autoregressive model, into pretraining. Empirically, under comparable experiment setting, XLNet outperforms BERT on 20 tasks, often by a large margin, including question answering, natural language inference, sentiment analysis, and document ranking.
Zhilin Yang 0001, Zihang Dai, Yiming Yang 0002, Jaime G. Carbonell, Ruslan Salakhutdinov, Quoc V. Le
NeurIPS4
2019 Toward Reciprocity-Aware Distributed Learning in Referral Networks
Ashiqur R. KhudaBukhsh, Jaime G. Carbonell
PRICAI (2)2
2019 Expertise drift in referral networks
Ashiqur R. KhudaBukhsh, Jaime G. Carbonell
Auton. Agents Multi Agent Syst.2
2018 Adapting Word Embeddings to New Languages with Morphological and Phonological Subword Representations
abstract
Much work in Natural Language Processing (NLP) has been for resource-rich languages, making generalization to new, less-resourced languages challenging.We present two approaches for improving generalization to lowresourced languages by adapting continuous word representations using linguistically motivated subword units: phonemes, morphemes and graphemes.Our method requires neither parallel corpora nor bilingual dictionaries and provides a significant gain in performance over previous methods relying on these resources.We demonstrate the effectiveness of our approaches on Named Entity Recognition for four languages, namely Uyghur, Turkish, Bengali and Hindi, of which Uyghur and Bengali are low resource languages, and also perform experiments on Machine Translation.Exploiting subwords with transfer learning gives us a boost of +15.2 NER F1 for Uyghur and +9.7 F1 for Bengali.We also show improvements in the monolingual setting where we achieve (avg.)+3 F1 and (avg.)+1.35 BLEU.
Aditi Chaudhary, Chunting Zhou, Lori S. Levin, Graham Neubig, David R. Mortensen, Jaime G. Carbonell
EMNLP6
2018 DeepCx: A transition-based approach for shallow semantic parsing with complex constructional triggers
abstract
This paper introduces the SURFACE CON-STRUCTION LABELING (SCL) task, which expands the coverage of Shallow Semantic Parsing (SSP) to include frames triggered by complex constructions.We present DeepCx, a neural, transition-based system for SCL.As a test case for the approach, we apply DeepCx to the task of tagging causal language in English, which relies on a wider variety of constructions than are typically addressed in SSP.We report substantial improvements over previous tagging efforts on a causal language dataset.We also propose ways DeepCx could be extended to still more difficult constructions and to other semantic domains once appropriate datasets become available.
Jesse Dunietz, Jaime G. Carbonell, Lori S. Levin
EMNLP2
2018 Towards Semi-Supervised Learning for Deep Semantic Role Labeling
abstract
Neural models have shown several state-ofthe-art performances on Semantic Role Labeling (SRL).However, the neural models require an immense amount of semantic-role corpora and are thus not well suited for lowresource languages or domains.The paper proposes a semi-supervised semantic role labeling method that outperforms the state-ofthe-art in limited SRL training corpora.The method is based on explicitly enforcing syntactic constraints by augmenting the training objective with a syntactic-inconsistency loss component and uses SRL-unlabeled instances to train a joint-objective LSTM.On CoNLL-2012 English section, the proposed semi-supervised training with 1%, 10% SRLlabeled data and varying amounts of SRLunlabeled data achieves +1.58, +0.78 F1, respectively, over the pre-trained models that were trained on SOTA architecture with ELMo on the same SRL-labeled data.Additionally, by using the syntactic-inconsistency loss on inference time, the proposed model achieves +3.67, +2.1 F1 over pre-trained model on 1%, 10% SRL-labeled data, respectively.
Sanket Vaibhav Mehta, Jay-Yoon Lee, Jaime G. Carbonell
EMNLP3
2018 Neural Cross-lingual Named Entity Recognition with Minimal Resources
abstract
For languages with no annotated resources, unsupervised transfer of natural language processing models such as named-entity recognition (NER) from resource-rich languages would be an appealing capability.However, differences in words and word order across languages make it a challenging problem.To improve mapping of lexical items across languages, we propose a method that finds translations based on bilingual word embeddings.To improve robustness to word order differences, we propose to use self-attention, which allows for a degree of flexibility with respect to word order.We demonstrate that these methods achieve state-of-the-art or competitive NER performance on commonly tested languages under a cross-lingual setting, with much lower resource requirements than past approaches.We also evaluate the challenges of applying these methods to Uyghur, a lowresource language.1
Jiateng Xie, Zhilin Yang 0001, Graham Neubig, Noah A. Smith, Jaime G. Carbonell
EMNLP5
2018 Temporal transfer learning for drift adaptation
Daegun Won, Peter J. Jansen, Jaime G. Carbonell
ESANN3
2018 Endorsement in Referral Networks
Ashiqur R. KhudaBukhsh, Jaime G. Carbonell
EUMAS2
2018 Market-Aware Proactive Skill Posting
Ashiqur R. KhudaBukhsh, Jong Woo Hong, Jaime G. Carbonell
ISMIS3
2018 Towards More Reliable Transfer Learning
Jaime G. Carbonell
ECML/PKDD (2)2
2018 Robust learning in expert networks: a comparative analysis
Ashiqur R. KhudaBukhsh, Jaime G. Carbonell, Peter J. Jansen
J. Intell. Inf. Syst.2
2018 Bounds on the minimax rate for estimating a prior over a VC class from independent learning tasks
Liu Yang 0001, Steve Hanneke, Jaime G. Carbonell
Theor. Comput. Sci.3
2017 Vision-Language Fusion for Object Recognition
abstract
While recent advances in computer vision have caused object recognition rates to spike, there is still much room for improvement. In this paper, we develop an algorithm to improve object recognition by integrating human-generated contextual information with vision algorithms. Specifically, we examine how interactive systems such as robots can utilize two types of context information--verbal descriptions of an environment and human-labeled datasets. We propose a re-ranking schema, MultiRank, for object recognition that can efficiently combine such information with the computer vision results. In our experiments, we achieve up to 9.4% and 16.6% accuracy improvements using the oracle and the detected bounding boxes, respectively, over the vision-only recognizers. We conclude that our algorithm has the ability to make a significant impact on object recognition in robotics and beyond.
Sz-Rung Shiang, Stephanie Rosenthal, Anatole Gershman, Jaime G. Carbonell, Jean Oh
AAAI4
2017 Nonparametric Neural Networks
George Philipp, Jaime G. Carbonell
ICLR (Poster)2
2017 Completely Heterogeneous Transfer Learning with Attention - What And What Not To Transfer
abstract
We study a transfer learning framework where source and target datasets are heterogeneous in both feature and label spaces. Specifically, we do not assume explicit relations between source and target tasks a priori, and thus it is crucial to determine what and what not to transfer from source knowledge. Towards this goal, we define a new heterogeneous transfer learning approach that (1) selects and attends to an optimized subset of source samples to transfer knowledge from, and (2) builds a unified transfer network that learns from both source and target knowledge. This method, termed "Attentional Heterogeneous Transfer", along with a newly proposed unsupervised transfer loss, improve upon the previous state-of-the-art approaches on extensive simulations as well as a challenging hetero-lingual text classification task.
Seungwhan Moon, Jaime G. Carbonell
IJCAI2
2017 Self-Paced Multitask Learning with Shared Knowledge
abstract
This paper introduces self-paced task selection to multitask learning, where instances from more closely related tasks are selected in a progression of easier-to-harder tasks, to emulate an effective human education strategy, but applied to multitask machine learning. We develop the mathematical foundation for the approach based on iterative selection of the most appropriate task, learning the task parameters, and updating the shared knowledge, optimizing a new bi-convex loss function. This proposed method applies quite generally, including to multitask feature learning, multitask learning with alternating structure optimization, etc. Results show that in each of the above formulations self-paced (easier-to-harder) task selection outperforms the baseline version of these methods in all the experiments.
Keerthiram Murugesan, Jaime G. Carbonell
IJCAI2
2017 Robust Learning in Expert Networks: A Comparative Analysis
Ashiqur R. KhudaBukhsh, Jaime G. Carbonell, Peter J. Jansen
ISMIS2
2017 Active Learning from Peers
abstract
This paper addresses the challenge of learning from peers in an online multitask setting. Instead of always requesting a label from a human oracle, the proposed method first determines if the learner for each task can acquire that label with sufficient confidence from its peers either as a task-similarity weighted sum, or from the single most similar task. If so, it saves the oracle query for later use in more difficult cases, and if not it queries the human oracle. The paper develops the new algorithm to exhibit this behavior and proves a theoretical mistake bound for the method compared to the best linear predictor in hindsight. Experiments over three multitask learning benchmark datasets show clearly superior performance over baselines such as assuming task independence, learning only from the oracle and not learning from peer tasks.
Keerthiram Murugesan, Jaime G. Carbonell
NIPS2
2017 Multi-Task Multiple Kernel Relationship Learning
abstract
This paper presents a novel multitask multiple-kernel learning framework that efficiently learns the kernel weights leveraging the relationship across multiple tasks. The idea is to automatically infer this task relationship in the RKHS space corresponding to the given base kernels. The problem is formulated as a regularization-based approach called Multi-Task Multiple Kernel Relationship Learning (MK-MTRL), which models the task relationship matrix from the weights learned from latent feature spaces of task-specific base kernels. Unlike in previous work, the proposed formulation allows one to incorporate prior knowledge for simultaneously learning several related task. We propose an alternating minimization algorithm to learn the model parameters, kernel weights and task relationship matrix. In order to tackle large-scale problems, we further propose a two-stage MK-MTRL online learning algorithm and show that it significantly reduces the computational time, and also achieves performance comparable to that of the joint learning framework. Experimental results on benchmark datasets show that the proposed formulations outperform several state-of-the-art multitask learning methods.
Keerthiram Murugesan, Jaime G. Carbonell
SDM2
2017 Event-based summarization using a centrality-as-relevance model
Luís Marujo, Ricardo Ribeiro 0001, Anatole Gershman, David Martins de Matos, João Paulo da Silva Neto, Jaime G. Carbonell
Knowl. Inf. Syst.6
2017 Automatically Tagging Constructions of Causation and Their Slot-Fillers
abstract
This paper explores extending shallow semantic parsing beyond lexical-unit triggers, using causal relations as a test case. Semantic parsing becomes difficult in the face of the wide variety of linguistic realizations that causation can take on. We therefore base our approach on the concept of constructions from the linguistic paradigm known as Construction Grammar (CxG). In CxG, a construction is a form/function pairing that can rely on arbitrary linguistic and semantic features. Rather than codifying all aspects of each construction’s form, as some attempts to employ CxG in NLP have done, we propose methods that offload that problem to machine learning. We describe two supervised approaches for tagging causal constructions and their arguments. Both approaches combine automatically induced pattern-matching rules with statistical classifiers that learn the subtler parameters of the constructions. Our results show that these approaches are promising: they significantly outperform naïve baselines for both construction recognition and cause and effect head matches.
Jesse Dunietz, Lori S. Levin, Jaime G. Carbonell
Trans. Assoc. Comput. Linguistics3
2016 Leveraging Multilingual Training for Limited Resource Event Extraction
abstract
Event extraction has become one of the most important topics in information extraction, but to date, there is very limited work on leveraging cross-lingual training to boost performance. We propose a new event extraction approach that trains on multiple languages using a combination of both language-dependent and language-independent features, with particular focus on the case where target domain training data is of very limited size. We show empirically that multilingual training can boost performance for the tasks of event trigger extraction and event argument extraction on the Chinese ACE 2005 dataset.
Andrew Hsi, Yiming Yang 0002, Jaime G. Carbonell, Ruochen Xu
COLING3
2016 Distributed Learning in Expert Referral Networks
abstract
Human experts or autonomous agents in a referral network must decide whether to accept a task or refer to a more appropriate expert, and if so to whom. In order for the referral network to improve over time, the experts must learn to estimate the topical expertise of other experts. This paper extends concepts from Reinforcement Learning and Active Learning to referral networks, to learn how to refer at the network level, based on the proposed distributed interval estimation learning (DIEL) algorithm. Diverse Monte Carlo simulations reveal that DIEL improves network performance significantly over both greedy and Q-learning baselines [3], approaching optimal given enough data.
Ashiqur R. KhudaBukhsh, Peter J. Jansen, Jaime G. Carbonell
ECAI3
2016 Data-driven Automated Induction of Prerequisite Structure Graphs
Devendra Singh Chaplot, Yiming Yang 0002, Jaime G. Carbonell, Kenneth R. Koedinger
EDM3
2016 Phonologically Aware Neural Model for Named Entity Recognition in Low Resource Transfer Settings
abstract
Named Entity Recognition is a well established information extraction task with many state of the art systems existing for a variety of languages.Most systems rely on language specific resources, large annotated corpora, gazetteers and feature engineering to perform well monolingually.In this paper, we introduce an attentional neural model which only uses language universal phonological character representations with word embeddings to achieve state of the art performance in a monolingual setting using supervision and which can quickly adapt to a new language with minimal or no data.We demonstrate that phonological character representations facilitate cross-lingual transfer, outperform orthographic representations and incorporating both attention and phonological features improves statistical efficiency of the model in 0-shot and low data transfer settings with no task specific feature engineering in the source or target language.
Akash Bharadwaj, David R. Mortensen, Chris Dyer, Jaime G. Carbonell
EMNLP4
2016 Efficient Shift-Invariant Dictionary Learning
abstract
Shift-invariant dictionary learning (SIDL) refers to the problem of discovering a set of latent basis vectors (the dictionary) that captures informative local patterns at different locations of the input sequences, and a sparse coding for each sequence as a linear combination of the latent basis elements. It differs from conventional dictionary learning and sparse coding where the latent basis has the same dimension as the input vectors, where the focus is on global patterns instead of shift-invariant local patterns. Unsupervised discovery of shift-invariant dictionary and the corresponding sparse coding has been an open challenge as the number of candidate local patterns is extremely large, and the number of possible linear combinations of such local patterns is even more so. In this paper we propose a new framework for unsupervised discovery of both the shift-invariant basis and the sparse coding of input data, with efficient algorithms for tractable optimization. Empirical evaluations on multiple time series data sets demonstrate the effectiveness and efficiency of the proposed method.
Guoqing Zheng, Yiming Yang 0002, Jaime G. Carbonell
KDD3
2016 Generation from Abstract Meaning Representation using Tree Transducers
abstract
Jeffrey Flanigan, Chris Dyer, Noah A. Smith, Jaime Carbonell. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Jeffrey Flanigan, Chris Dyer, Noah A. Smith, Jaime G. Carbonell
HLT-NAACL4
2016 Adaptive Smoothed Online Multi-Task Learning
abstract
This paper addresses the challenge of jointly learning both the per-task model parameters and the inter-task relationships in a multi-task online learning setting. The proposed algorithm features probabilistic interpretation, efficient updating rules and flexible modulation on whether learners focus on their specific task or on jointly address all tasks. The paper also proves a sub-linear regret bound as compared to the best linear predictor in hindsight. Experiments over three multi-task learning benchmark datasets show advantageous performance of the proposed approach over several state-of-the-art online multi-task learning baselines.
Keerthiram Murugesan, Hanxiao Liu, Jaime G. Carbonell, Yiming Yang 0002
NIPS3
2016 Proactive Transfer Learning for Heterogeneous Feature and Label Spaces
Seungwhan Moon, Jaime G. Carbonell
ECML/PKDD (2)2
2016 Multitask Matrix Completion for Learning Protein Interactions Across Diseases
Meghana Kshirsagar 0001, Jaime G. Carbonell, Judith Klein-Seetharaman, Keerthiram Murugesan
RECOMB2
2016 Learning Concept Graphs from Online Educational Data
abstract
This paper addresses an open challenge in educational data mining, i.e., the problem of automatically mapping online courses from different providers (universities, MOOCs, etc.) onto a universal space of concepts, and predicting latent prerequisite dependencies (directed links) among both concepts and courses. We propose a novel approach for inference within and across course-level and concept-level directed graphs. In the training phase, our system projects partially observed course-level prerequisite links onto directed concept-level links; in the testing phase, the induced concept-level links are used to infer the unknown course-level prerequisite links. Whereas courses may be specific to one institution, concepts are shared across different providers. The bi-directional mappings enable our system to perform interlingua-style transfer learning, e.g. treating the concept graph as the interlingua and transferring the prerequisite relations across universities via the interlingua. Experiments on our newly collected datasets of courses from MIT, Caltech, Princeton and CMU show promising results.
Hanxiao Liu, Yiming Yang 0002, Jaime G. Carbonell
J. Artif. Intell. Res.4
2016 Exploring events and distributed representations of text in multi-document summarization
Luís Marujo, Wang Ling, Ricardo Ribeiro 0001, Anatole Gershman, Jaime G. Carbonell, David Martins de Matos, João Paulo da Silva Neto
Knowl. Based Syst.5
2015 Unsupervised Phrasal Near-Synonym Generation from Text Corpora
abstract
Unsupervised discovery of synonymous phrases is useful in a variety of tasks ranging from text mining and search engines to semantic analysis and machine translation. This paper presents an unsupervised corpus-based conditional model: Near-Synonym System (NeSS) for finding phrasal synonyms and near synonyms that requires only a large monolingual corpus. The method is based on maximizing information-theoretic combinations of shared contexts and is parallelizable for large-scale processing. An evaluation framework with crowd-sourced judgments is proposed and results are compared with alternate methods, demonstrating considerably superior results to the literature and to thesaurus look up for multi-word phrases. Moreover, the results show that the statistical scoring functions and overall scalability of the system are more important than language specific NLP tools. The method is language-independent and practically useable due to accuracy and real-time performance via parallel decomposition.
Dishan Gupta, Jaime G. Carbonell, Anatole Gershman, Steve Klein
AAAI2
2015 Bounds on the Minimax Rate for Estimating a Prior over a VC Class from Independent Learning Tasks
Liu Yang 0001, Steve Hanneke, Jaime G. Carbonell
ALT3
2015 Concept Graph Learning from Educational Data
abstract
This paper addresses an open challenge in educational data mining, i.e., the problem of using observed prerequisite relations among courses to learn a directed universal concept graph, and using the induced graph to predict unobserved prerequisite relations among a broader range of courses. This is particularly useful to induce prerequisite relations among courses from different providers (universities, MOOCs, etc.). We propose a new framework for inference within and across two graphs---at the course level and at the induced concept level---which we call Concept Graph Learning (CGL). In the training phase, our system projects the course-level links onto the concept space to induce directed concept links; in the testing phase, the concept links are used to predict (unobserved) prerequisite links for test-set courses within the same institution or across institutions. The dual mappings enable our system to perform an interlingua-style transfer learning, e.g. treating the concept graph as the interlingua, and inducing prerequisite links in a transferable manner across different universities. Experiments on our newly collected data sets of courses from MIT, Caltech, Princeton and CMU show promising results, including the viability of CGL for transfer learning.
Yiming Yang 0002, Hanxiao Liu, Jaime G. Carbonell
WSDM3
2014 A Discriminative Graph-Based Parser for the Abstract Meaning Representation
abstract
Meaning Representation (AMR) is a semantic formalism for which a growing set of annotated examples is available.We introduce the first approach to parse sentences into this representation, providing a strong baseline for future improvement.The method is based on a novel algorithm for finding a maximum spanning, connected subgraph, embedded within a Lagrangian relaxation of an optimization problem that imposes linguistically inspired constraints.Our approach is described in the general framework of structured prediction, allowing future incorporation of additional features and constraints, and may extend to other formalisms as well.Our open-source system, JAMR, is available at:
Jeffrey Flanigan, Sam Thomson, Jaime G. Carbonell, Chris Dyer, Noah A. Smith
ACL (1)3
2014 Proactive learning with multiple class-sensitive labelers
abstract
Proactive learning extends active learning by considering multiple labelers with different accuracies and costs, thus optimizing labeler selection as well as instance selection. In this paper, we propose a novel method to estimate labeler accuracy per class and to select labelers based on both cost and estimated accuracy, combined with an ensemble approach called multi-class information density (MCID) as a selection criterion. Our approach relaxes the common assumption found in past work that labeler accuracy is independent of class for multi-class learning, and by estimating the class-conditional accuracy better assigns instances to labelers. Results on several datasets with both real and simulated experts strongly demonstrate the efficacy of these methods.
Seungwhan Moon, Jaime G. Carbonell
DSAA2
2014 Detecting Non-Adversarial Collusion in Crowdsourcing
abstract
A group of agents are said to collude if they share information or make joint decisions in a manner contrary to explicit or implicit social rules that results in an unfair advantage over non-colluding agents or other interested parties. For instance, collusion manifests as sharing answers in exams, as colluding bidders in auctions, or as colluding participants (e.g., Turkers) in crowd sourcing. This paper studies the latter, where the goal of the colluding participants is to "earn" money without doing the actual work, for instance by copying product ratings of another colluding participant, adding limited noise as attempted obfuscation. Such collusion not only yields fewer independent ratings, but may also introduce strong biases in aggregate results if undetected. Our proposed unsupervised collusion detection algorithm identifies colluding groups in crowd sourcing with fairly high accuracy both in synthetic and real data, and results in significant bias reduction, such as minimizing shifts from the true mean in rating tasks and recovering the true variance among raters.
Ashiqur R. KhudaBukhsh, Jaime G. Carbonell, Peter J. Jansen
HCOMP2
2014 Saddle Points and Accelerated Perceptron Algorithms
abstract
In this paper, we consider the problem of finding a linear (binary) classifier or providing a near-infeasibility certificate if there is none. We bring a new perspective to addressing these two problems simultaneously in a single efficient process, by investigating a related Bilinear Saddle Point Problem (BSPP). More specifically, we show that a BSPP-based approach provides either a linear classifier or an ε-infeasibility certificate. We show that the accelerated primal-dual algorithm, Mirror Prox, can be used for this purpose and achieves the best known convergence rate of O(\sqrt\log n\overρ(A)) (O(\sqrt\log n\overε)), which is \emphalmost independent of the problem size, n. Our framework also solves kernelized and conic versions of the problem, with the same rate of convergence. We support our theoretical findings with an empirical study on synthetic and real data, highlighting the efficiency and numerical stability of our algorithms, especially on large-scale instances.
Adams Wei Yu, Fatma Kilinç-Karzan, Jaime G. Carbonell
ICML3
2014 Resources for the Detection of Conventionalized Metaphors in Four Languages
Lori S. Levin, Teruko Mitamura, Brian MacWhinney, Davida Fromm, Jaime G. Carbonell, Weston Feely, Robert E. Frederking, Anatole Gershman
LREC5
2014 Efficient Structured Matrix Rank Minimization
Adams Wei Yu, Yaoliang Yu, Jaime G. Carbonell, Suvrit Sra
NIPS4
2013 Learnability of DNF with representation-specific queries
abstract
We study the problem of PAC learning the class of DNF formulas with a type of natural pairwise query specific to the DNF representation. Specifically, given a pair of positive examples from a polynomial-sized sample, we consider boolean queries that ask whether the two examples satisfy at least one term in common in the target DNF, and numerical queries that ask how many terms in common the two examples satisfy. We provide both positive and negative results for learning with these queries under both uniform and general distributions.
Liu Yang 0001, Avrim Blum, Jaime G. Carbonell
ITCS3
2013 Large-Scale Discriminative Training for Statistical Machine Translation Using Held-Out Line Search
Jeffrey Flanigan, Chris Dyer, Jaime G. Carbonell
HLT-NAACL3
2013 Buy-in-Bulk Active Learning
abstract
In many practical applications of active learning, it is more cost-effective to request labels in large batches, rather than one-at-a-time. This is because the cost of labeling a large batch of examples at once is often sublinear in the number of examples in the batch. In this work, we study the label complexity of active learning algorithms that request labels in a given number of batches, as well as the tradeoff between the total number of queries and the number of rounds allowed. We additionally study the total cost sufficient for learning, for an abstract notion of the cost of requesting the labels of a given number of examples at once. In particular, we find that for sublinear cost functions, it is often desirable to request labels in large batches (i.e., buying in bulk); although this may increase the total number of labels requested, it reduces the total cost required for learning.
Liu Yang 0001, Jaime G. Carbonell
NIPS2
2013 Self reinforcement for important passage retrieval
abstract
In general, centrality-based retrieval models treat all elements of the retrieval space equally, which may reduce their effectiveness. In the specific context of extractive summarization (or important passage retrieval), this means that these models do not take into account that information sources often contain lateral issues, which are hardly as important as the description of the main topic, or are composed by mixtures of topics. We present a new two-stage method that starts by extracting a collection of key phrases that will be used to help centrality-as-relevance retrieval model. We explore several approaches to the integration of the key phrases in the centrality model. The proposed method is evaluated using different datasets that vary in noise (noisy vs clean) and language (Portuguese vs English). Results show that the best variant achieves relative performance improvements of about 31% in clean data and 18% in noisy data.
Ricardo Ribeiro 0001, Luís Marujo, David Martins de Matos, João Paulo da Silva Neto, Anatole Gershman, Jaime G. Carbonell
SIGIR6
2013 Multitask learning for host-pathogen protein interactions
abstract
MOTIVATION: An important aspect of infectious disease research involves understanding the differences and commonalities in the infection mechanisms underlying various diseases. Systems biology-based approaches study infectious diseases by analyzing the interactions between the host species and the pathogen organisms. This work aims to combine the knowledge from experimental studies of host-pathogen interactions in several diseases to build stronger predictive models. Our approach is based on a formalism from machine learning called 'multitask learning', which considers the problem of building models across tasks that are related to each other. A 'task' in our scenario is the set of host-pathogen protein interactions involved in one disease. To integrate interactions from several tasks (i.e. diseases), our method exploits the similarity in the infection process across the diseases. In particular, we use the biological hypothesis that similar pathogens target the same critical biological processes in the host, in defining a common structure across the tasks. RESULTS: Our current work on host-pathogen protein interaction prediction focuses on human as the host, and four bacterial species as pathogens. The multitask learning technique we develop uses a task-based regularization approach. We find that the resulting optimization problem is a difference of convex (DC) functions. To optimize, we implement a Convex-Concave procedure-based algorithm. We compare our integrative approach to baseline methods that build models on a single host-pathogen protein interaction dataset. Our results show that our approach outperforms the baselines on the training data. We further analyze the protein interaction predictions generated by the models, and find some interesting insights. AVAILABILITY: The predictions and code are available at: http://www.cs.cmu.edu/∼mkshirsa/ismb2013_paper320.html . SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Meghana Kshirsagar 0001, Jaime G. Carbonell, Judith Klein-Seetharaman
Bioinform.2
2013 A theory of transfer learning with applications to active learning
Liu Yang 0001, Steve Hanneke, Jaime G. Carbonell
Mach. Learn.3
2012 Collaborative workflow for crowdsourcing translation
abstract
In this paper we explore the challenges in crowdsourcing the task of translation over the web in which remotely located translators work on providing translations independent of each other. We then propose a collaborative workflow for crowdsourcing translation to address some of these challenges. In our pipeline model, the translators are working in phases where output from earlier phases can be enhanced in the subsequent phases. We also highlight some of the novel contributions of the pipeline model like assistive translation and translation synthesis that can leverage monolingual and bilingual speakers alike. We evaluate our approach by eliciting translations for both a minority-to-majority language pair and a minority-to-minority language pair. We observe that in both scenarios, our workflow produces better quality translations in a cost-effective manner, when compared to the traditional crowdsourcing workflow.
Vamshi Ambati, Stephan Vogel, Jaime G. Carbonell
CSCW3
2012 Cost-Sensitive Risk Stratification in the Diagnosis of Heart Disease
abstract
We investigate machine learning methods for diagnostic screening of heart disease. Coronary heart disease is the leading cause of death in the US, causing more deaths than all types of cancers combined. Early diagnosis of heart disease in women is harder than it is in men and typically requires the administration of several clinical tests on the patient. Most risk stratification methods aggregate the results of such tests, including the risky, invasive procedures that cannot be administered on all patients. In this paper, our goal is to identify patients who are under high-risk of having heart disease and related adverse events, using a minimal number of diagnostic tests, especially less invasive ones. The low frequency of patients with severe heart disease in the dataset is challenging for most conventional machine learning methods. To overcome this problem, we develop and apply a cost-sensitive k nearest neighbor algorithm. Our contributions are two fold: First, we compare the predictive value of several diagnostic procedures for heart disease, including electrocardiography, angiography, radionuclide perfusion and conclude that in womens heart disease, certain combinations of noninvasive techniques are more predictive than some of the widely used invasive procedures. Then, we evaluate held out data and achieve an AUROC over 0.70, signifying valuable clinical utility, using only the least costly and least invasive tests.
Selen Uguroglu, Robert Biederman, Jaime G. Carbonell
IAAI4
2012 Supervised Topical Key Phrase Extraction of News Stories using Crowdsourcing, Light Filtering and Co-reference Normalization
Luís Marujo, Anatole Gershman, Jaime G. Carbonell, Robert E. Frederking, João Paulo da Silva Neto
LREC3
2012 Adaptive Multi-task Sparse Learning with an Application to fMRI Study
abstract
In this paper, we consider the multi-task sparse learning problem under the assumption that the dimensionality diverges with the sample size. The traditional l1/l2 multi-task lasso does not enjoy the oracle property unless a rather strong condition is enforced. Inspired by adaptive lasso, we propose a multi-stage procedure, adaptive multi-task lasso, to simultaneously conduct model estimation and variable selection across different tasks. Motivated by adaptive elastic-net, we further propose the adaptive multi-task elastic-net by adding another quadratic penalty to address the problem of collinearity. When the number of tasks is fixed, under weak assumptions, we establish the asymptotic oracle property for the proposed adaptive multi-task sparse learning methods including both adaptive multi-task lasso and elastic-net. In addition to the desirable asymptotic property, we show by simulations that adaptive sparse learning methods also achieve much improved finite sample performance. As a case study, we apply adaptive multi-task elastic-net to a cognitive science problem, where one wants to discover a compact semantic basis for predicting fMRI images. We show that adaptive multi-task sparse learning methods achieve superior performance and provide some insights into how the brain represents meanings of words.
Xi Chen 0010, Jingrui He, Rick Lawrence, Jaime G. Carbonell
SDM4
2012 Techniques to cope with missing data in host-pathogen protein interaction prediction
abstract
MOTIVATION: Approaches that use supervised machine learning techniques for protein-protein interaction (PPI) prediction typically use features obtained by integrating several sources of data. Often certain attributes of the data are not available, resulting in missing values. In particular, our host-pathogen PPI datasets have a large fraction, in the range of 58-85% of missing values, which makes it challenging to apply machine learning algorithms. RESULTS: We show that specialized techniques for missing value imputation can improve the performance of the models significantly. We use cross species information in combination with machine learning techniques like Group lasso with ℓ(1)/ℓ(2) regularization. We demonstrate the benefits of our approach on two PPI prediction problems. In our first example of Salmonella-human PPI prediction, we are able to obtain high prediction accuracies with 77.6% precision and 84% recall. Comparison with various other techniques shows an improvement of 9 in F1 score over the next best technique. We also apply our method to Yersinia-human PPI prediction successfully, demonstrating the generality of our approach. AVAILABILITY: Predicted interactions, datasets, features are available at: http://www.cs.cmu.edu/~mkshirsa/eccb2012_paper46.html. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Meghana Kshirsagar 0001, Jaime G. Carbonell, Judith Klein-Seetharaman
Bioinform.2
2012 An effective framework for characterizing rare categories
Jingrui He, Hanghang Tong, Jaime G. Carbonell
Frontiers Comput. Sci.3
2011 Phrasal Equivalence Classes for Generalized Corpus-Based Machine Translation
Rashmi Gangadharaiah, Ralf D. Brown, Jaime G. Carbonell
CICLing (2)3
2011 Modeling personalized email prioritization: classification-based and regression-based approaches
abstract
Email overload, even after spam filtering, presents a serious productivity challenge for busy professionals and executives. One solution is automated prioritization of incoming emails to ensure the most important are read and processed quickly, while others are processed later as/if time permits in declining priority levels. This paper presents a study of machine learning approaches to email prioritization into discrete levels, comparing ordinal regression versus classifier cascades. Given the ordinal nature of discrete email priority levels, SVM ordinal regression would be expected to perform well, but surprisingly a cascade of SVM classifiers significantly outperforms ordinal regression for email prioritization. In contrast, SVM regression performs well -- better than classifiers -- on selected UCI data sets. This unexpected performance inversion is analyzed and results are presented, providing core functionality for email prioritization systems.
Shinjae Yoo, Yiming Yang 0002, Jaime G. Carbonell
CIKM3
2011 Multi-Strategy Approaches to Active Learning for Statistical Machine Translation
Vamshi Ambati, Stephan Vogel, Jaime G. Carbonell
MTSummit3
2011 Feature Selection for Transfer Learning
Selen Uguroglu, Jaime G. Carbonell
ECML/PKDD (3)2
2011 Sparse Latent Semantic Analysis
abstract
Latent semantic analysis (LSA), as one of the most popular unsupervised dimension reduction tools, has a wide range of applications in text mining and information retrieval. The key idea of LSA is to learn a projection matrix that maps the high dimensional vector space representations of documents to a lower dimensional latent space, i.e. so called latent topic space. In this paper, we propose a new model called Sparse LSA, which produces a sparse projection matrix via the ℓ1 regularization. Compared to the traditional LSA, Sparse LSA selects only a small number of relevant words for each topic and hence provides a compact representation of topic-word relationships. Moreover, Sparse LSA is computationally very efficient with much less memory usage for storing the projection matrix. Furthermore, we propose two important extensions of Sparse LSA: group structured Sparse LSA and non-negative Sparse LSA. We conduct experiments on several benchmark datasets and compare Sparse LSA and its extensions with several widely used methods, e.g. LSA, Sparse Coding and LDA. Empirical results suggest that Sparse LSA achieves similar performance gains to LSA, but is more efficient in projection computation, storage, and also well explain the topic-word relationships.
Xi Chen 0010, Yanjun Qi, Qihang Lin, Jaime G. Carbonell
SDM5
2011 SmartNotes: Application of crowdsourcing to the detection of web threats
abstract
We describe a crowdsourcing system, called SmartNotes, which detects security threats related to web browsing, such as Internet scams, deceptive sales of substandard products, and websites with intentionally misleading information. It combines automatically collected data about websites with user votes and comments, and uses them to identify potential threats. We have implemented it as a browser extension, which is available for free public use.
Mehrbod Sharifi, Eugene Fink, Jaime G. Carbonell
SMC3
2011 Detection of Internet scam using logistic regression
abstract
Internet scam is fraudulent or intentionally misleading information posted on the web, usually with the intent of tricking people into sending money or disclosing sensitive information. We describe an application of logistic regression to detection of Internet scam. The developed system automatically collects 43 characteristic statistics about websites from 11 online sources and computes the probability that a given website is malicious. We present its empirical evaluation, which shows that its precision and recall are about 98%.
Mehrbod Sharifi, Eugene Fink, Jaime G. Carbonell
SMC3
2011 Smoothing Proximal Gradient Method for General Structured Sparse Learning
Xi Chen 0010, Qihang Lin, Jaime G. Carbonell, Eric P. Xing
UAI4
2010 Learning Spatial-Temporal Varying Graphs with Applications to Climate Data Analysis
abstract
An important challenge in understanding climate change is to uncover the dependency relationships between various climate observations and forcing factors. Graphical lasso, a recently proposed L1 penalty based structure learning algorithm, has been proven successful for learning underlying dependency structures for the data drawn from a multivariate Gaussian distribution. However, climatological data often turn out to be non-Gaussian, e.g. cloud cover, precipitation, etc. In this paper, we examine nonparametric learning methods to address this challenge. In particular, we develop a methodology to learn dynamic graph structures from spatial-temporal data so that the graph structures at adjacent time or locations are similar. Experimental results demonstrate that our method not only recovers the underlying graph well but also captures the smooth variation properties on both synthetic data and climate data. An important challenge in understanding climate change is to uncover the dependency relationships between various climate observations and forcing factors. Graphical lasso, a recently proposed An important challenge in understanding climate change is to uncover the dependency relationships between various climate observations and forcing factors. Graphical lasso, a recently proposed L1 penalty based structure learning algorithm, has been proven successful for learning underlying dependency structures for the data drawn from a multivariate Gaussian distribution. However, climatological data often turn out to be non-Gaussian, e.g. cloud cover, precipitation, etc. In this paper, we examine nonparametric learning methods to address this challenge. In particular, we develop a methodology to learn dynamic graph structures from spatial-temporal data so that the graph structures at adjacent time or locations are similar. Experimental results demonstrate that our method not only recovers the underlying graph well but also captures the smooth variation properties on both synthetic data and climate data.
Xi Chen 0010, Yan Liu 0002, Han Liu 0001, Jaime G. Carbonell
AAAI4
2010 Bayesian Active Learning Using Arbitrary Binary Valued Queries
Liu Yang 0001, Steve Hanneke, Jaime G. Carbonell
ALT3
2010 Rank learning for factoid question answering with linguistic and semantic constraints
abstract
This work presents a general rank-learning framework for passage ranking within Question Answering (QA) systems using linguistic and semantic features. The framework enables query-time checking of complex linguistic and semantic constraints over keywords. Constraints are composed of a mixture of keyword and named entity features, as well as features derived from semantic role labeling. The framework supports the checking of constraints of arbitrary length relating any number of keywords. We show that a trained ranking model using this rich feature set achieves greater than a 20% improvement in Mean Average Precision over baseline keyword retrieval models. We also show that constraints based on semantic role labeling features are particularly effective for passage retrieval; when they can be leveraged, an 40% improvement in MAP over the baseline can be realized.
Matthew W. Bilotti, Jonathan L. Elsas, Jaime G. Carbonell, Eric Nyberg
CIKM3
2010 Automatic Determination of Number of clusters for creating Templates in Example-Based Machine Translation
Rashmi Gangadharaiah, Ralf D. Brown, Jaime G. Carbonell
EAMT3
2010 Chunk-Based EBMT
Jae Dong Kim, Ralf D. Brown, Jaime G. Carbonell
EAMT3
2010 Learning Preferences with Millions of Parameters by Enforcing Sparsity
abstract
We study the retrieval task that ranks a set of objects for a given query in the pair wise preference learning framework. Recently researchers found out that raw features (e.g. words for text retrieval) and their pair wise features which describe relationships between two raw features (e.g. word synonymy or polysemy) could greatly improve the retrieval precision. However, most existing methods can not scale up to problems with many raw features (e.g. English vocabulary), due to the prohibitive computational cost on learning and the memory requirement to store a quadratic number of parameters. In this paper, we propose to learn a sparse representation of the pair wise features under the preference learning framework using the L1 regularization. Based on stochastic gradient descent, an online algorithm is devised to enforce the sparsity using a mini-batch shrinkage strategy. On multiple benchmark datasets, we show that our method achieves better performance with fast convergence, and takes much less memory on models with millions of parameters.
Xi Chen 0010, Yanjun Qi, Qihang Lin, Jaime G. Carbonell
ICDM5
2010 Rare Category Characterization
abstract
Rare categories abound and their characterization has heretofore received little attention. Fraudulent banking transactions, network intrusions, and rare diseases are examples of rare classes whose detection and characterization are of high value. However, accurate characterization is challenging due to high-skewness and non-separability from majority classes, e.g., fraudulent transactions masquerade as legitimate ones. This paper proposes the RACH algorithm by exploring the compactness property of the rare categories. It is based on an optimization framework which encloses the rare examples by a minimum-radius hyper ball. The framework is then converted into a convex optimization problem, which is in turn effectively solved in its dual form by the projected sub gradient method. RACH can be naturally kernelized. Experimental results validate the effectiveness of RACH.
Jingrui He, Hanghang Tong, Jaime G. Carbonell
ICDM3
2010 Automatic Detection of HIV Drug Resistance-Associated Mutations
abstract
Each HIV-1 patient has a diverse population of virus strains in his/her body as the virus quickly replicates and mutates, requiring a combination drug therapy optimized to the patient's unique viral population. Towards this goal, prediction systems have been developed to deduce the susceptibility of a given HIV genotype to a single drug. Many are rule-based systems or rely on hand-crafted features which are difficult to update for HIV strains and new drugs. We adapted the vector-of-n-grams approach from document classification and chi-square feature selection to automatically generate a feature set that yields comparable performance to the expert-selected and database-derived feature sets without requiring treatment history data. Our automatically-generated feature set also found all the expert-selected mutations and more demonstrating its potential for knowledge discovery. Compared to the previous state-of-the-art with ample expert knowledge, our best fully-automated prediction model for each drug yielded comparable performance at 82.9% classification accuracy and 0.819 coefficient of determination on average. Along with its lack of need for human expertise and potential for knowledge discovery, our automatic feature selection method is a good candidate for the more complex prediction task of combination drug therapy optimization.
Betty Yee Man Cheng, Jaime G. Carbonell
ICMLA2
2010 Active Learning and Crowd-Sourcing for Machine Translation
Vamshi Ambati, Stephan Vogel, Jaime G. Carbonell
LREC3
2010 A Probabilistic Framework to Learn from Multiple Annotators with Time-Varying Accuracy
abstract
This paper addresses the challenging problem of learning from multiple annotators whose labeling accuracy (reliability) differs and varies over time. We propose a framework based on Sequential Bayesian Estimation to learn the expected accuracy at each time step while simultaneously deciding which annotators to query for a label in an incremental learning framework. We develop a variant of the particle filtering method that estimates the expected accuracy at every time step by sets of weighted samples and performs sequential Bayes updates. The estimated expected accuracies are then used to decide which annotators to be queried at the next time step. The empirical analysis shows that the proposed method is very effective at predicting the true label using only moderate labeling efforts, resulting in cleaner labels to train classifiers. The proposed method significantly outperforms a repeated labeling baseline which queries all labelers per example and takes the majority vote to predict the true label. Moreover, our method is able to track the true accuracy of an annotator quite well in the absence of gold standard labels. These results demonstrate the strength of the proposed method in terms of estimating the time-varying reliability of multiple annotators and producing cleaner, better quality labels without extensive label queries.
Pinar Donmez, Jaime G. Carbonell, Jeff G. Schneider
SDM2
2010 Co-selection of Features and Instances for Unsupervised Rare Category Analysis
abstract
Rare category analysis is of key importance both in theory and in practice. Previous research work focuses on supervised rare category analysis, such as rare category detection and rare category classification. In this paper, for the first time, we address the challenge of unsupervised rare category analysis, including feature selection and rare category selection. We propose to jointly deal with the two correlated tasks so that they can benefit from each other. To this end, we design an optimization framework which is able to co-select the relevant features and the examples from the rare category (a.k.a. the minority class). It is well justified theoretically. Furthermore, we develop the Partial Augmented Lagrangian Method (PALM) to solve the optimization problem. Experimental results on both synthetic and real data sets show the effectiveness of the proposed method.
Jingrui He, Jaime G. Carbonell
SDM2
2010 Temporal Collaborative Filtering with Bayesian Probabilistic Tensor Factorization
abstract
Real-world relational data are seldom stationary, yet traditional collaborative filtering algorithms generally rely on this assumption. Motivated by our sales prediction problem, we propose a factor-based algorithm that is able to take time into account. By introducing additional factors for time, we formalize this problem as a tensor factorization with a special constraint on the time dimension. Further, we provide a fully Bayesian treatment to avoid tuning parameters and achieve automatic model complexity control. To learn the model we develop an efficient sampling procedure that is capable of analyzing large-scale data sets. This new algorithm, called Bayesian Probabilistic Tensor Factorization (BPTF), is evaluated on several real-world problems including sales prediction and movie recommendation. Empirical results demonstrate the superiority of our temporal model.
Liang Xiong, Xi Chen 0010, Tzu-Kuo Huang, Jeff G. Schneider, Jaime G. Carbonell
SDM5
2010 Learning of personalized security settings
abstract
While many cybersecurity tools are available to computer users, their default configurations often do not match needs of specific users. Since most modern users are not computer experts, they are often unable to customize these tools, thus getting either insufficient or excessive security. To address this problem, we are developing an automated assistant that learns security needs of the user and helps customize available tools.
Mehrbod Sharifi, Eugene Fink, Jaime G. Carbonell
SMC3
2010 Semi-supervised multi-task learning for predicting interactions between HIV-1 and human proteins
abstract
MOTIVATION: Protein-protein interactions (PPIs) are critical for virtually every biological function. Recently, researchers suggested to use supervised learning for the task of classifying pairs of proteins as interacting or not. However, its performance is largely restricted by the availability of truly interacting proteins (labeled). Meanwhile, there exists a considerable amount of protein pairs where an association appears between two partners, but not enough experimental evidence to support it as a direct interaction (partially labeled). RESULTS: We propose a semi-supervised multi-task framework for predicting PPIs from not only labeled, but also partially labeled reference sets. The basic idea is to perform multi-task learning on a supervised classification task and a semi-supervised auxiliary task. The supervised classifier trains a multi-layer perceptron network for PPI predictions from labeled examples. The semi-supervised auxiliary task shares network layers of the supervised classifier and trains with partially labeled examples. Semi-supervision could be utilized in multiple ways. We tried three approaches in this article, (i) classification (to distinguish partial positives with negatives); (ii) ranking (to rate partial positive more likely than negatives); (iii) embedding (to make data clusters get similar labels). We applied this framework to improve the identification of interacting pairs between HIV-1 and human proteins. Our method improved upon the state-of-the-art method for this task indicating the benefits of semi-supervised multi-task learning using auxiliary information. AVAILABILITY: http://www.cs.cmu.edu/~qyj/HIVsemi.
Yanjun Qi, Öznur Tastan, Jaime G. Carbonell, Judith Klein-Seetharaman, Jason Weston
Bioinform.3
2010 Active learning for human protein-protein interaction prediction
abstract
BACKGROUND: Biological processes in cells are carried out by means of protein-protein interactions. Determining whether a pair of proteins interacts by wet-lab experiments is resource-intensive; only about 38,000 interactions, out of a few hundred thousand expected interactions, are known today. Active machine learning can guide the selection of pairs of proteins for future experimental characterization in order to accelerate accurate prediction of the human protein interactome. RESULTS: Random forest (RF) has previously been shown to be effective for predicting protein-protein interactions. Here, four different active learning algorithms have been devised for selection of protein pairs to be used to train the RF. With labels of as few as 500 protein-pairs selected using any of the four active learning methods described here, the classifier achieved a higher F-score (harmonic mean of Precision and Recall) than with 3000 randomly chosen protein-pairs. F-score of predicted interactions is shown to increase by about 15% with active learning in comparison to that with random selection of data. CONCLUSION: Active learning algorithms enable learning more accurate classifiers with much lesser labelled data and prove to be useful in applications where manual annotation of data is formidable. Active learning techniques demonstrated here can also be applied to other proteomics applications such as protein structure prediction and classification.
Thahir P. Mohamed, Jaime G. Carbonell, Madhavi Ganapathiraju
BMC Bioinform.2
2010 Active machine learning for transmembrane helix prediction
abstract
BACKGROUND: About 30% of genes code for membrane proteins, which are involved in a wide variety of crucial biological functions. Despite their importance, experimentally determined structures correspond to only about 1.7% of protein structures deposited in the Protein Data Bank due to the difficulty in crystallizing membrane proteins. Algorithms that can identify proteins whose high-resolution structure can aid in predicting the structure of many previously unresolved proteins are therefore of potentially high value. Active machine learning is a supervised machine learning approach which is suitable for this domain where there are a large number of sequences but only very few have known corresponding structures. In essence, active learning seeks to identify proteins whose structure, if revealed experimentally, is maximally predictive of others. RESULTS: An active learning approach is presented for selection of a minimal set of proteins whose structures can aid in the determination of transmembrane helices for the remaining proteins. TMpro, an algorithm for high accuracy TM helix prediction we previously developed, is coupled with active learning. We show that with a well-designed selection procedure, high accuracy can be achieved with only few proteins. TMpro, trained with a single protein achieved an F-score of 94% on benchmark evaluation and 91% on MPtopo dataset, which correspond to the state-of-the-art accuracies on TM helix prediction that are achieved usually by training with over 100 training proteins. CONCLUSION: Active learning is suitable for bioinformatics applications, where manually characterized data are not a comprehensive representation of all possible data, and in fact can be a very sparse subset thereof. It aids in selection of data instances which when characterized experimentally can improve the accuracy of computational characterization of remaining raw data. The results presented here also demonstrate that the feature extraction method of TMpro is well designed, achieving a very good separation between TM and non TM segments.
Hatice U. Osmanbeyoglu, Jessica A. Wehner, Jaime G. Carbonell, Madhavi Ganapathiraju
BMC Bioinform.3
2009 Active Sampling for Rank Learning via Optimizing the Area under the ROC Curve
Pinar Donmez, Jaime G. Carbonell
ECIR2
2009 Accelerated Gradient Method for Multi-task Sparse Learning Problem
abstract
Many real world learning problems can be recast as multi-task learning problems which utilize correlations among different tasks to obtain better generalization performance than learning each task individually. The feature selection problem in multi-task setting has many applications in fields of computer vision, text classification and bio-informatics. Generally, it can be realized by solving a L-1-infinity regularized optimization problem. And the solution automatically yields the joint sparsity among different tasks. However, due to the nonsmooth nature of the L-1-infinity norm, there lacks an efficient training algorithm for solving such problem with general convex loss functions. In this paper, we propose an accelerated gradient method based on an ``optimal'' first order black-box method named after Nesterov and provide the convergence rate for smooth convex loss functions. For nonsmooth convex loss functions, such as hinge loss, our method still has fast convergence rate empirically. Moreover, by exploiting the structure of the L-1-infinity ball, we solve the black-box oracle in Nesterov's method by a simple sorting scheme. Our method is suitable for large-scale multi-task learning problem since it only utilizes the first order information and is very easy to implement. Experimental results show that our method significantly outperforms the most state-of-the-art methods in both convergence speed and learning accuracy.
Xi Chen 0010, Weike Pan, James T. Kwok, Jaime G. Carbonell
ICDM4
2009 Efficiently learning the accuracy of labeling sources for selective sampling
abstract
Many scalable data mining tasks rely on active learning to provide the most useful accurately labeled instances. However, what if there are multiple labeling sources ('oracles' or 'experts') with different but unknown reliabilities? With the recent advent of inexpensive and scalable online annotation tools, such as Amazon's Mechanical Turk, the labeling process has become more vulnerable to noise - and without prior knowledge of the accuracy of each individual labeler. This paper addresses exactly such a challenge: how to jointly learn the accuracy of labeling sources and obtain the most informative labels for the active learning task at hand minimizing total labeling effort. More specifically, we present IEThresh (Interval Estimate Threshold) as a strategy to intelligently select the expert(s) with the highest estimated labeling accuracy. IEThresh estimates a confidence interval for the reliability of each expert and filters out the one(s) whose estimated upper-bound confidence interval is below a threshold - which jointly optimizes expected accuracy (mean) and need to better estimate the expert's accuracy (variance). Our framework is flexible enough to work with a wide range of different noise levels and outperforms baselines such as asking all available experts and random expert selection. In particular, IEThresh achieves a given level of accuracy with less than half the queries issued by all-experts labeling and less than a third the queries required by random expert selection on datasets such as the UCI mushroom one. The results show that our method naturally balances exploration and exploitation as it gains knowledge of which experts to rely upon, and selects them with increasing frequency.
Pinar Donmez, Jaime G. Carbonell, Jeff G. Schneider
KDD2
2009 Prior-Free Rare Category Detection
abstract
Rare category detection is an open challenge in machine learning. It plays the central role in applications such as detecting new financial fraud patterns, detecting new network malware, and scientific discovery. In such cases rare categories are hidden among huge volumes of normal data and observations. In this paper, we propose a new method for rare category detection named SEDER, which requires no prior information about the data set. It implicitly performs semiparametric density estimation using specially designed exponentially families, and then picks the examples for labeling where the neighborhood density changes the most. SEDER can work in the cases where the data is not separable. Its unique feature over all existing methods lies in its prior-free nature, i.e. it does not require any prior information about the data set (e.g. the number of classes, the proportion of the different classes, etc.). Therefore, it is more suitable for real applications. Experimental results on both synthetic and real data sets demonstrate the superiority of SEDER.
Jingrui He, Jaime G. Carbonell
SDM2
2009 It pays to be picky: an evaluation of thread retrieval in online forums
abstract
Online forums host a rich information exchange, often with contributions from many subject matter experts. In this work we evaluate algorithms for thread retrieval in a large and active online forum community. We compare methods that utilize thread structure to a naïve method that treats a thread as a single document. We find that thread structure helps, and additionally selective methods of thread scoring, which only use evidence from a small number of messages in the thread, significantly and consistently outperform inclusive methods which use all the messages in the thread.
Jonathan L. Elsas, Jaime G. Carbonell
SIGIR2
2009 Scheduling with uncertain resources: Representation of common knowledge
abstract
We describe a system for scheduling a conference based on incomplete information about available resources and scheduling constraints. We explain the representation of uncertain knowledge and related common-sense rules, which allow reasoning based on uncertain and partially missing data.
Eugene Fink, Matt Jennings, Konstantin Salomatin, Jaime G. Carbonell
SMC4
2009 Analysis of uncertain data: Smoothing of histograms
abstract
We consider the problem of converting a set of numeric data points into a smoothed approximation of the underlying probability distribution. We describe a representation of distributions by histograms with variable-width bars, and give a greedy smoothing algorithm based on this representation.
Eugene Fink, Ankur Sarin, Jaime G. Carbonell
SMC3
2009 Creating and visualizing fuzzy document classification
abstract
Fuzzy classification ranks items by degree rather than assigning them either within or without of a category. The novelty of our work is in integrating fuzzy classification algorithms with an interface to visualize fuzzy results. An advantage of our algorithms' `fuzziness' is that it provides additional information per retrieved result that helps in deciding whether to drill down to the document or skip it. An advantage of our interface is that it allows users to visualize those differences quickly. We have created a prototype that allows the retrieval of journal articles by content word or by ontology-supported browse categories that can be selected independently or in tandem. Journal articles in our digital library pertain to paleontology, but techniques demonstrated viable in indexing and ranking paleo-journal literature should apply to other knowledge domains with little modification.
Judith Gelernter, Raymond Lu, Eugene Fink, Jaime G. Carbonell
SMC5
2009 Analysis of uncertain data: Selection of probes for information gathering
abstract
We consider the problem of gathering data for evaluation of given hypotheses, and describe a method for analyzing tradeoffs between the expected utility and the cost of data collection.
Anatole Gershman, Eugene Fink, Jaime G. Carbonell
SMC4
2009 Analysis of uncertain data: Evaluation of given hypotheses
abstract
We consider the problem of heuristic evaluation of given hypotheses based on limited observations, in situations when available data are insufficient for rigorous statistical analysis.
Anatole Gershman, Eugene Fink, Jaime G. Carbonell
SMC4
2008 RADAR: A Personal Assistant that Learns to Reduce Email Overload
Michael Freed, Jaime G. Carbonell, Geoffrey J. Gordon, Jordan Hayes, Brad A. Myers, Daniel P. Siewiorek, Stephen F. Smith, Aaron Steinfeld, Anthony Tomasic
AAAI2
2008 Suppressing outliers in pairwise preference ranking
abstract
Many of the recently proposed algorithms for learning feature-based ranking functions are based on the pairwise preference framework, in which instead of taking documents in isolation, document pairs are used as instances in the learning process. One disadvantage of this process is that a noisy relevance judgment on a single document can lead to a large number of mis-labeled document pairs. This can jeopardize robustness and deteriorate overall ranking performance. In this paper we study the effects of outlying pairs in rank learning with pairwise preferences and introduce a new meta-learning algorithm capable of suppressing these undesirable effects. This algorithm works as a second optimization step in which any linear baseline ranker can be used as input. Experiments on eight different ranking datasets show that this optimization step produces statistically significant performance gains over state-of-the-art methods.
Vitor R. Carvalho, Jonathan L. Elsas, William W. Cohen, Jaime G. Carbonell
CIKM4
2008 Proactive learning: cost-sensitive active learning with multiple imperfect oracles
abstract
Proactive learning is a generalization of active learning designed to relax unrealistic assumptions and thereby reach practical applications. Active learning seeks to select the most informative unlabeled instances and ask an omniscient oracle for their labels, so as to retrain the learning algorithm maximizing accuracy. However, the oracle is assumed to be infallible (never wrong), indefatigable (always answers), individual (only one oracle), and insensitive to costs (always free or always charges the same). Proactive learning relaxes all four of these assumptions, relying on a decision-theoretic approach to jointly select the optimal oracle and instance, by casting the problem as a utility optimization problem subject to a budget constraint. Results on multi-oracle optimization over several data sets demonstrate the superiority of our approach over the single-imperfect-oracle baselines in most cases.
Pinar Donmez, Jaime G. Carbonell
CIKM2
2008 Corpus microsurgery: criteria optimization for medical cross-language ir
abstract
Automatic subset selection from a parallel corpus significantly cross-lingual information retrieval (CLIR) performance, in addition to increasing its efficiency. Our selection method extracts relevant training data by incorporating additional criteria (i.e. estimated corpus quality, taxonomy projection and size) in addition to lexical-based criteria. The challenge lies in combining these criteria using a meaningful scoring function that can be used for ranking parallel sentence candidates. We choose weighted geometric mean for its soft-AND properties, and we optimize criteria weights by wrapping the CLIR task in an optimization shell. Due to the indeterminate nature of the search space convexity properties, we have explored continuous reactive tabu search (CRTS), a global optimization method. We use a large parallel corpus in the medical domain to examine the effect of adaptation criteria and their combination on CLIR performance. In our experiments, 100 selected sentences yield 90% of the performance obtained with 5,000 times more in-domain parallel sentences. Our optimized criteria weights considerably outperform the uniform distribution baseline, as well as lexical
Monica Rogati, Yiming Yang 0002, Jaime G. Carbonell
CIKM3
2008 Optimizing estimated loss reduction for active sampling in rank learning
abstract
Learning to rank is becoming an increasingly popular research area in machine learning. The ranking problem aims to induce an ordering or preference relations among a set of instances in the input space. However, collecting labeled data is growing into a burden in many rank applications since labeling requires eliciting the relative ordering over the set of alternatives. In this paper, we propose a novel active learning framework for SVM-based and boosting-based rank learning. Our approach suggests sampling based on maximizing the estimated loss differential over unlabeled data. Experimental results on two benchmark corpora show that the proposed model substantially reduces the labeling effort, and achieves superior performance rapidly with as much as 30% relative improvement over the margin-based sampling baseline.
Pinar Donmez, Jaime G. Carbonell
ICML2
2008 Document Representation and Query Expansion Models for Blog Recommendation
Jaime Arguello, Jonathan L. Elsas, Jamie Callan, Jaime G. Carbonell
ICWSM4
2008 Cluster-Based Query Expansion for Statistical Question Answering
Lucian Vlad Lita, Jaime G. Carbonell
IJCNLP2
2008 Predicate Indexing for Incremental Multi-Query Optimization
Chun Jin, Jaime G. Carbonell
ISMIS2
2008 Linguistic Structure and Bilingual Informants Help Induce Machine Translation of Lesser-Resourced Languages
Christian Monson, Ariadna Font Llitjós, Vamshi Ambati, Lori S. Levin, Alon Lavie, Alison Alvarez, Roberto Aranovich, Jaime G. Carbonell, Robert E. Frederking, Erik Peterson, Katharina Probst
LREC8
2008 Retrieval and feedback models for blog feed search
abstract
Blog feed search poses different and interesting challenges from traditional ad hoc document retrieval. The units of retrieval, the blogs, are collections of documents, the blog posts. In this work we adapt a state-of-the-art federated search model to the feed retrieval task, showing a significant improvement over algorithms based on the best performing submissions in the TREC 2007 Blog Distillation task[12]. We also show that typical query expansion techniques such as pseudo-relevance feedback using the blog corpus do not provide any significant performance improvement and in many cases dramatically hurt performance. We perform an in-depth analysis of the behavior of pseudo-relevance feedback for this task and develop a novel query expansion technique using the link structure in Wikipedia. This query expansion technique provides significant and consistent performance improvements for this task, yielding a 22% and 14% improvement in MAP over the unexpanded query for our baseline and federated algorithms respectively.
Jonathan L. Elsas, Jaime Arguello, Jamie Callan, Jaime G. Carbonell
SIGIR4
2008 The impact of history length on personalized search
abstract
Personalized search is a promising way to better serve different users' information needs. Search history is one of the major information sources for search personalization. We investigated the impact of history length on the effectiveness of personalized ranking. We carried out task-based user study for Web search, and obtained ranked relevance judgments for all queries. Query contexts derived from previous queries in the same task are used to re-rank results for the current query. Experimental results show that the performance of personalization generally improves as more queries are accumulated, but most of the benefits come from a few immediately preceding queries.
Yangbo Zhu, Jamie Callan, Jaime G. Carbonell
SIGIR3
2008 Scheduling with uncertain resources: Learning to ask the right questions
abstract
We consider the task of scheduling a conference based on incomplete information about resources and constraints, which requires elicitation of additional data, and describe a learning procedure that improves elicitation strategies. We outline the representation of incomplete knowledge, and then describe an adaptive elicitation procedure, which learns to identify critical missing data.
Alexander Carpentier, Mehrbod Sharifi, Eugene Fink, Jaime G. Carbonell
SMC4
2008 Analysis of uncertain data: Tools for representation and processing
abstract
We present initial work on a general-purpose system for the analysis of incomplete and uncertain data, integrated with Excel. We explain the representation of main types of uncertainty, and outline tools for the analysis of uncertain data and planning of additional data collection.
Eugene Fink, Jaime G. Carbonell
SMC3
2008 Scheduling with uncertain resources: Learning to make reasonable assumptions
abstract
We consider the task of scheduling a conference based on incomplete information about resources and constraints, and describe a mechanism for the dynamic learning of related default assumptions, which enable the scheduling system to make reasonable guesses about missing data. We outline the representation of incomplete knowledge, describe the learning procedure, and demonstrate that the learned knowledge improves the scheduling results.
Steven Gardiner, Eugene Fink, Jaime G. Carbonell
SMC3
2008 Fast learning of document ranking functions with the committee perceptron
abstract
This paper presents a new variant of the perceptron algorithm using selective committee averaging (or voting). We apply this agorithm to the problem of learning ranking functions for document retrieval, known as the "Learning to Rank" problem. Most previous algorithms proposed to address this problem focus on minimizing the number of misranked document pairs in the training set. The committee perceptron algorithm improves upon existing solutions by biasing the final solution towards maximizing an arbitrary rank-based performance metrics. This method performs comparably or better than two state-of-the-art rank learning algorithms, and also provides significant training time improvements over those methods, showing over a 45-fold reduction in training time compared to ranking SVM
Jonathan L. Elsas, Vitor R. Carvalho, Jaime G. Carbonell
WSDM3
2007 Combining N-grams and Alignment in G-protein Coupling Specificity Prediction
Betty Yee Man Cheng, Jaime G. Carbonell
APBC2
2007 Genre identification and goal-focused summarization
abstract
In this paper, we present a novel technique of first performing document genre identification, then utilizing the genre for producing tailored summaries based on a user's information seeking needs - genre oriented goal-focused summarization - such as a plot or opinion summary of a movie review. We create a test corpus to determine genre classification accuracy for 16 genres, and examine performance on various amounts of training data for machine learning algorithms - Random Forests, SVM light and Naïve Bayes. Results show that Random Forests outperforms SVM light and Naïve Bayes. The genre tag is used to inform a downstream summarization engine. We define types of summaries for 7 genres, create a ground truth corpus and analyze the results of genre oriented goal-focused summarization, showing that this type of user based summarization requires different algorithms than the leading sentence baseline which is known to perform well in the case of news articles.
Jade Goldstein-Stewart, Gary M. Ciany, Jaime G. Carbonell
CIKM3
2007 Dual Strategy Active Learning
Pinar Donmez, Jaime G. Carbonell, Paul N. Bennett
ECML2
2007 Graph-Based Semi-Supervised Learning as a Generative Model
Jingrui He, Jaime G. Carbonell, Yan Liu 0002
IJCAI2
2007 Cluster-Based Selection of Statistical Answering Strategies
Lucian Vlad Lita, Jaime G. Carbonell
IJCAI2
2007 Protein Quaternary Fold Recognition Using Conditional Graphical Models
Yan Liu 0002, Jaime G. Carbonell, Vanathi Gopalakrishnan, Peter Weigele
IJCAI2
2007 Improving transfer-based MT systems with automatic refinements
Ariadna Font Llitjós, Jaime G. Carbonell, Alon Lavie
MTSummit2
2007 Combining Probability-Based Rankers for Action-Item Detection
Paul N. Bennett, Jaime G. Carbonell
HLT-NAACL2
2007 Nearest-Neighbor-Based Active Learning for Rare Category Detection
abstract
Rare category detection is an open challenge for active learning, especially in the de-novo case (no labeled examples), but of significant practical importance for data mining - e.g. detecting new financial transaction fraud patterns, where normal legitimate transactions dominate. This paper develops a new method for detecting an instance of each minority class via an unsupervised local-density-differential sampling strategy. Essentially a variable-scale nearest neighbor process is used to optimize the probability of sampling tightly-grouped minority classes, subject to a local smoothness assumption of the majority class. Results on both synthetic and real data sets are very positive, detecting each minority class with only a frac- tion of the actively sampled points required by random sampling and by Pelleg’s Interleave method, the prior best technique in the sparse literature on this topic.
Jingrui He, Jaime G. Carbonell
NIPS2
2006 ARGUS: Efficient Scalable Continuous Query Optimization for Large-Volume Data Streams
abstract
We present the architecture of ARGUS, a stream processing system implemented atop commercial DBMSs to support large-scale complex continuous queries over data streams. ARGUS supports incremental operator evaluation and incremental multi-query plan optimization as new queries arrive. The latter is done to a degree well beyond the previous state-of-the-art via a suite of techniques such as query-algebra canonicalization, indexing, and searching, and topological query network optimization with join order optimization, conditional materialization, minimal column projection, and transitivity inference. Building on top of a DBMS, the system provides a value-adding package to the existing database applications where the needs of stream processing become increasingly demanding. Compared to directly running the continuous queries on the DBMS, ARGUS achieves well over a 100-fold improvement in performance
Chun Jin, Jaime G. Carbonell
IDEAS2
2006 Incremental Aggregation on Multiple Continuous Queries
Chun Jin, Jaime G. Carbonell
ISMIS2
2006 Spectral Clustering for Example Based Machine Translation
Rashmi Gangadharaiah, Ralf D. Brown, Jaime G. Carbonell
HLT-NAACL3
2006 Scheduling with Uncertain Resources: Representation and Utility Function
abstract
We describe the representation of uncertain knowledge in a conference-scheduling system, which may include incomplete information about available resources, conference events, and scheduling constraints. We then explain the use of this incomplete knowledge in the evaluation of schedule quality.
Ulas Bardak, Eugene Fink, Jaime G. Carbonell
SMC3
2006 Scheduling with Uncertain Resources: Elicitation of Additional Data
abstract
We consider the task of scheduling a conference based on incomplete data about available resource and scheduling constraints, and describe a procedure for automated elicitation of additional data. This procedure is part of an interactive system for scheduling under uncertainty, which identifies critical missing information, generates related questions to the human administrator, and uses answers to improve the schedule.
Ulas Bardak, Eugene Fink, Chris R. Martens, Jaime G. Carbonell
SMC4
2006 Scheduling with Uncertain Resources: Collaboration with the User
abstract
We describe a scheduling system that supports collaboration between the user and automated optimizer. It enables the user to monitor the optimizer decisions, make any of the decisions manually, and leave the other decisions to the system. Furthermore, it identifies the tasks that require the user's participation, and asks for assistance with these tasks.
Eugene Fink, Ulas Bardak, Brandon Rothrock, Jaime G. Carbonell
SMC4
2006 Scheduling with Uncertain Resources: Search for a Near-Optimal Solution
abstract
We describe a system for scheduling a conference based on incomplete information about available resources and scheduling constraints. We explain the representation of uncertain knowledge, describe a local-search algorithm for generating near-optimal schedules, and give empirical results of automated scheduling under uncertainty.
Eugene Fink, P. Matthew Jennings, Ulas Bardak, Jean Oh, Stephen F. Smith, Jaime G. Carbonell
SMC6
2005 Symmetric probabilistic alignment for example-based translation
Jae Dong Kim, Ralf D. Brown, Peter J. Jansen, Jaime G. Carbonell
EAMT4
2005 A framework for interactive and automatic refinement of transfer-based machine translation
Ariadna Font Llitjós, Jaime G. Carbonell, Alon Lavie
EAMT2
2005 Predicting protein folds with structural repeats using a chain graph model
abstract
Protein fold recognition is a key step towards inferring the tertiary structures from amino-acid sequences. Complex folds such as those consisting of interacting structural repeats are prevalent in proteins involved in a wide spectrum of biological functions. However, extant approaches often perform inadequately due to their inability to capture long-range interactions between structural units and to handle low sequence similarities across proteins (under 25% identity). In this paper, we propose a chain graph model built on a causally connected series of segmentation conditional random fields (SCRFs) to address these issues. Specifically, the SCRF model captures long-range interactions within recurring structural units and the Bayesian network backbone decomposes cross-repeat interactions into locally computable modules consisting of repeat-specific SCRFs and a model for sequence motifs. We applied this model to predict β -helices and leucine-rich repeats, and found it significantly outperforms extant methods in predictive accuracy and/or computational efficiency.
Yan Liu 0002, Eric P. Xing, Jaime G. Carbonell
ICML3
2005 A Machine Text-Inspired Machine Learning Approach for Identification of Transmembrane Helix Boundaries
Betty Yee Man Cheng, Jaime G. Carbonell, Judith Klein-Seetharaman
ISMIS2
2005 ARGUS: Rete + DBMS = Efficient Persistent Profile Matching on Large-Volume Data Streams
Chun Jin, Jaime G. Carbonell, Philip J. Hayes
ISMIS2
2005 Segmentation Conditional Random Fields (SCRFs): A New Approach for Protein Fold Recognition
Yan Liu 0002, Jaime G. Carbonell, Peter Weigele, Vanathi Gopalakrishnan
RECOMB2
2005 Detecting action-items in e-mail
abstract
No abstract available.
Paul N. Bennett, Jaime G. Carbonell
SIGIR2
2004 Unsupervised question answering data acquisition from local corpora
abstract
Data-driven approaches in question answering (QA) are increasingly common. Since availability of training data for such approaches is very limited, we propose an unsupervised algorithm that generates high quality question-answer pairs from local corpora. The algorithm is ontology independent, requiring very small seed data as its starting point. Two alternating views of the data make learning possible: 1) question types are viewed as relations between entities and 2) question types are described by their corresponding question-answer pairs. These two aspects of the data allow us to construct an unsupervised algorithm that acquires high precision question-answer pairs. We show the quality of the acquired data for different question types and perform a task-based evaluation. With each iteration, pairs acquired by the unsupervised algorithm are used as training data to a simple QA system. Performance increases with the number of question-answer pairs acquired confirming the robustness of the unsupervised algorithm. We introduce the notion of semantic drift and show that it is a desirable quality in training data for question answering systems.
Lucian Vlad Lita, Jaime G. Carbonell
CIKM2
2004 Instance-Based Question Answering: A Data-Driven Approach
Lucian Vlad Lita, Jaime G. Carbonell
EMNLP2
2004 Developing Language Resources for a Transnational Digital Government System
Violetta Cavalli-Sforza, Jaime G. Carbonell, Peter J. Jansen
LREC2
2004 The Translation Correction Tool: English-Spanish User Studies
Ariadna Font Llitjós, Jaime G. Carbonell
LREC2
2004 Data Collection and Analysis of Mapudungun Morphology for Spelling Correction
Christian Monson, Lori S. Levin, Rodolfo Vega, Ralf D. Brown, Ariadna Font Llitjós, Alon Lavie, Jaime G. Carbonell, Eliseo Cañulef, Rosendo Huisca
LREC7
2004 Context sensitive vocabulary and its application in protein secondary structure prediction
abstract
Protein secondary structure prediction is an important step towards understanding the relation between protein sequence and structure. However, most current prediction methods use features difficult for biologists to interpret. In this paper, we present a new method that applies information retrieval techniques to solve the problem:we extract a context sensitive biological vocabulary for protein sequences and apply text classification methods to predict protein secondary structure. Experimental results show that our method performs comparably to the state-of-art methods. Furthermore, the context sensitive vocabularies can serve as a useful tool to discover meaningful regular expression patterns for protein structures.
Yan Liu 0002, Jaime G. Carbonell, Judith Klein-Seetharaman, Vanathi Gopalakrishnan
SIGIR2
2004 Comparison of probabilistic combination methods for protein secondary structure prediction
abstract
MOTIVATION: Protein secondary structure prediction is an important step towards understanding how proteins fold in three dimensions. Recent analysis by information theory indicates that the correlation between neighboring secondary structures are much stronger than that of neighboring amino acids. In this article, we focus on the combination problem for sequences, i.e. combining the scores or assignments from single or multiple prediction systems under the constraint of a whole sequence, as a target for improvement in protein secondary structure prediction. RESULTS: We apply several graphical chain models to solve the combination problem and show that they are consistently more effective than the traditional window-based methods. In particular, conditional random fields (CRFs) moderately improve the predictions for helices and, more importantly, for beta sheets, which are the major bottleneck for protein secondary structure prediction.
Yan Liu 0002, Jaime G. Carbonell, Judith Klein-Seetharaman, Vanathi Gopalakrishnan
Bioinform.2
2003 Grand challenges for information management
abstract
No abstract available.
Jaime G. Carbonell
CIKM1
2003 A New Pairwise Ensemble Approach for Text Classification
Yan Liu 0002, Jaime G. Carbonell, Rong Jin 0001
ECML2
2003 Reducing boundary friction using translation-fragment overlap
abstract
Many corpus-based Machine Translation (MT) systems generate a number of partial translations which are then pieced together rather than immediately producing one overall translation. While this makes them more robust to ill-formed input, they are subject to disfluencies at phrasal translation boundaries even for well-formed input. We address this “boundary friction” problem by introducing a method that exploits overlapping phrasal translations and the increased confidence in translation accuracy they imply. We specify an efficient algorithm for producing translations using overlap. Finally, our empirical analysis indicates that this approach produces higher quality translations than the standard method of combining non-overlapping fragments generated by our Example-Based MT (EBMT) system in a peak-to-peak comparison.
Ralf D. Brown, Rebecca Hutchinson, Paul N. Bennett, Jaime G. Carbonell, Peter J. Jansen
MTSummit4
2003 Experiments with a Hindi-to-English transfer-based MT system under a miserly data scenario
abstract
We describe an experiment designed to evaluate the capabilities of our trainable transfer-based (Xfer) machine translation approach, as applied to the task of Hindi-to-English translation, and trained under an extremely limited data scenario. We compare the performance of the Xfer approach with two corpus-based approaches---Statistical MT (SMT) and Example-based MT (EBMT)---under the limited data scenario. The results indicate that the Xfer system significantly outperforms both EBMT and SMT in this scenario. Results also indicate that automatically learned transfer rules are effective in improving translation performance, compared with a baseline word-to-word translation version of the system. Xfer system performance with a limited number of manually written transfer rules is, however, still better than the current automatically inferred rules. Furthermore, a "multiengine" version of our system that combined the output of the Xfer and SMT systems and optimizes translation selection outperformed both individual systems.
Alon Lavie, Stephan Vogel, Lori S. Levin, Erik Peterson, Katharina Probst, Ariadna Font Llitjós, Rachel Reynolds, Jaime G. Carbonell, Richard Cohen
ACM Trans. Asian Lang. Inf. Process.8
2002 Boosting to correct inductive bias in text classification
abstract
This paper studies the effects of boosting in the context of different classification methods for text categorization, including Decision Trees, Naive Bayes, Support Vector Machines (SVMs) and a Rocchio-style classifier. We identify the inductive biases of each classifier and explore how boosting, as an error-driven resampling mechanism, reacts to those biases. Our experiments on the Reuters-21578 benchmark show that boosting is not effective in improving the performance of the base classifiers on common categories. However, the effect of boosting for rare categories varies across classifiers: for SVMs and Decision Trees, we achieved a 13-17% performance improvement in macro-averaged F1 measure, but did not obtain substantial improvement for the other two classifiers. This interesting finding of boosting on rare categories has not been reported before.
Yan Liu 0002, Yiming Yang 0002, Jaime G. Carbonell
CIKM3
2002 Topic-conditioned novelty detection
abstract
Automated detection of the first document reporting each new event in temporally-sequenced streams of documents is an open challenge. In this paper we propose a new approach which addresses this problem in two stages: 1) using a supervised learning algorithm to classify the on-line document stream into pre-defined broad topic categories, and 2) performing topic-conditioned novelty detection for documents in each topic. We also focus on exploiting named-entities for event-level novelty detection and using feature-based heuristics derived from the topic histories. Evaluating these methods using a set of broadcast news stories, our results show substantial performance gains over the traditional one-level approach to the novelty detection problem.
Yiming Yang 0002, Jian Zhang 0003, Jaime G. Carbonell, Chun Jin
KDD3
2002 MT for Minority Languages Using Elicitation-Based Learning of Syntactic Transfer Rules
Katharina Probst, Lori S. Levin, Erik Peterson, Alon Lavie, Jaime G. Carbonell
Mach. Transl.5
2000 Creating and Evaluating Multi-Document Sentence Extract Summaries
abstract
This paper discusses passage extraction approaches to multidocument summarization that use available information about the document set as a whole and the relationships between the documents to build on single document summarization methodology.Multi-document summarization diers from single in that the issues of compression, speed, redundancy and passage selection are critical in the formation of useful summaries, as well as the user's goals in creating the summary.Our approach addresses these issues by using domain-independent techniques based mainly on fast, statistical processing, a metric for reducing redundancy and maximizing diversity in the selected passages, and a modular framework to allow easy parameterization for dierent genres, corpora characteristics and user requirements.We examined how h umans create multi-document summaries as well as the characteristics of such summaries and use these summaries to evaluate the performance of various multidocument summarization algorithms.
Jade Goldstein-Stewart, Vibhu O. Mittal, Jaime G. Carbonell, Jamie Callan
CIKM3
2000 Special Issue of Machine Learning on Information Retrieval - Introduction
Jaime G. Carbonell, Yiming Yang 0002, William W. Cohen
Mach. Learn.1
1999 Summarizing Text Documents: Sentence Selection and Evaluation Metrics
abstract
Article Free Access Share on Summarizing text documents: sentence selection and evaluation metrics Authors: Jade Goldstein Language Technologies Institute, Carnegie Mellon University, Pittsburgh, PA Language Technologies Institute, Carnegie Mellon University, Pittsburgh, PAView Profile , Mark Kantrowitz Just Research, 4616 Henry Street, Pittsburgh, PA Just Research, 4616 Henry Street, Pittsburgh, PAView Profile , Vibhu Mittal Just Research, 4616 Henry Street, Pittsburgh, PA Just Research, 4616 Henry Street, Pittsburgh, PAView Profile , Jaime Carbonell Language Technologies Institute, Carnegie Mellon University, Pittsburgh, PA Language Technologies Institute, Carnegie Mellon University, Pittsburgh, PAView Profile Authors Info & Claims SIGIR '99: Proceedings of the 22nd annual international ACM SIGIR conference on Research and development in information retrievalAugust 1999 Pages 121–128https://doi.org/10.1145/312624.312665Published:01 August 1999Publication History 264citation3,037DownloadsMetricsTotal Citations264Total Downloads3,037Last 12 Months222Last 6 weeks42 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Jade Goldstein-Stewart, Mark Kantrowitz, Vibhu O. Mittal, Jaime G. Carbonell
SIGIR4
1998 The Use of MMR, Diversity-Based Reranking for Reordering Documents and Producing Summaries
abstract
No abstract available.
Jaime G. Carbonell, Jade Goldstein-Stewart
SIGIR1
1998 A Study of Retrospective and On-Line Event Detection
abstract
This paper investigates the use and extension of text retrieval and clustering techniques for event detection. The task is to automatically detect novel events from a temporally-ordered stream of news stories, either retrospectively or as the stories arrive. We applied hierarchical and non-hierarchical document clustering algorithms to a corpus of 15,836 stories, focusing on the exploitation of both content and temporal information. We found the resulting cluster hierarchies highly informative for retrospective detection of previously unidentified events, effectively supporting both query-free and query-driven retrieval. We also found that temporal distribution patterns of document clusters provide useful information for improvement in both retrospective detection and on-line detection of novel events. In an evaluation using manually labelled events to judge the system-detected events, we obtained a result of 82% in the Fl measure for retrospective detection, and a Fl value of 42% for on-line detection.
Yiming Yang 0002, Thomas Pierce, Jaime G. Carbonell
SIGIR3
1998 Translingual Information Retrieval: Learning from Bilingual Corpora
Yiming Yang 0002, Jaime G. Carbonell, Ralf D. Brown, Robert E. Frederking
Artif. Intell.2
1997 Translingual Information Retrieval: A Comparative Evaluation
Jaime G. Carbonell, Yiming Yang 0002, Robert E. Frederking, Ralf D. Brown, Yibing Geng, Danny Lee
IJCAI (1)1
1995 VERY Large Knowledge Bases - Architecture vs Engineering
James A. Hendler, Jaime G. Carbonell, Douglas B. Lenat, Riichiro Mizoguchi, Paul S. Rosenbloom
IJCAI2
1995 Integrating planning and learning: the PRODIGY architecture
abstract
Planning is a complex reasoning task that is well suited for the study of improving performance and knowledge by learning, i.e. by accumulation and interpretation of planning experience. PRODIGY is an architecture that integrates planning with multiple learning mechanisms. Learning occurs at the planner's decision points and integration in PRODIGY is achieved via mutually interpretable knowledge structures. This article describes the PRODIGY planner, briefly reports on several learning modules developed earlier along the project, and presents in more detail two recently explored methods to learn to generate plans of better quality. We introduce the techniques, illustrate them with comprehensive examples, and show preliminary empirical results. The article also includes a retrospective discussion of the characteristics of the overall PRODIGY architecture and discusses their evolution within the goal of the project of building a large and robust integrated planning and learning system.
Manuela M. Veloso, Jaime G. Carbonell, M. Alicia Pérez, Daniel Borrajo, Eugene Fink, Jim Blythe
J. Exp. Theor. Artif. Intell.2
1994 Evaluation Metrics for Knowledge-Based Machine Translation
Eric Nyberg, Teruko Mitamura, Jaime G. Carbonell
COLING3
1994 Knowledge Representation Issues in Integrated Planning and Learning Systems (Abstract)
Jaime G. Carbonell
KR1
1993 Lessons from TIPSTER/SHOGUN/JANUS
Jaime G. Carbonell
SIGIR1
1993 Derivational Analogy in Prodigy: Automating Case Acquisition, Storage, and Utilization
Manuela M. Veloso, Jaime G. Carbonell
Mach. Learn.2
1992 Is Production System Match Interesting?
abstract
A panel session in which issues relating to the effects of advances in faster and more parallel hardware, production system match (PSM) algorithms, and application domains for match on PSM as a research area is presented. It is argued that there is no such thing as the optimal matching algorithm, even for the well-defined task of production-system match and that broadening the scope of the matching task beyond forward-chaining production system presents a new set of problems to the artificial intelligence community. Also, even with all the speedups, large production system runs take hours to complete, and a major portion of this time is attributable to PSM. Match technology remains a large and centralized component of system performance. To that extent, providing sufficient speedups in the match in these systems may still be useful. Performance issues of production system execution are discussed, and a common set of benchmarks and test cases is called for. It is argued that parallel algorithms for match, resolve, and fire are all interesting and difficult problems to solve, and should be the focus of research by the PSM community.>
Mark W. Perlin, Jaime G. Carbonell, Daniel P. Miranker, Salvatore J. Stolfo, Milind Tambe
ICTAI2
1992 Machine Learning: A Maturing Field
Jaime G. Carbonell
Mach. Learn.1
1991 Editorial
Jaime G. Carbonell
Mach. Learn.1
1991 Editorial
Jaime G. Carbonell
Mach. Learn.1
1989 Towards a General Framework for Composing Disjunctive and Iterative Macro-operators
Peter Shell, Jaime G. Carbonell
IJCAI2
1989 Introduction: Paradigms for Machine Learning
Jaime G. Carbonell
Artif. Intell.1
1989 Explanation-Based Learning: A Problem Solving Perspective
Steven Minton, Jaime G. Carbonell, Craig A. Knoblock, Daniel Kuokka, Oren Etzioni, Yolanda Gil
Artif. Intell.2
1989 Editorial
Jaime G. Carbonell
Mach. Learn.1
1989 Editorial
Jaime G. Carbonell
Mach. Learn.1
1988 Anaphora resolution: a multy-strategy approach
Jaime G. Carbonell, Ralf D. Brown
COLING1
1987 Strategies for Learning Search Control Rules: An Explanation-based Approach
Steven Minton, Jaime G. Carbonell
IJCAI2
1987 The Universal Parser Architecture for Knowledge-based Machine Translation
Masaru Tomita, Jaime G. Carbonell
IJCAI2
1987 Integrating discourse pragmatics and propositional knowledge for multilingual natural language processing
Sergei Nirenburg, Jaime G. Carbonell
Mach. Transl.2
1986 The FERMI System: Inducing Iterative Macro-Operators from Experience
Patricia Cheng, Jaime G. Carbonell
AAAI2
1986 Parsing Spoken Language: A Semantic Caseframe Approach
Philip J. Hayes, Alex Hauptmann 0001, Jaime G. Carbonell, Masaru Tomita
COLING3
1986 Another Stride Towards Knowledge-Based Machine Translation
Masaru Tomita, Jaime G. Carbonell
COLING2
1986 Natural Language Interfaces - Ready for Commercial Success?
Wolfgang Wahlster, Jaime G. Carbonell, Gary G. Hendrix, Harry R. Tennant
COLING2
1986 Analogical reasoning in planning and decision making
abstract
Analogical reasoning is a significant cognitive process that has heretofore not been modelled by artificial intelligence researchers in a computationally tractable manner. Recently, several new approaches have shown significant promise using analogical processes for problem solving and learning. Among the first and most comprehensive, the ARIES project demonstrated that analogical problem solving is a computationally tractable means of exploiting past experience to solve new problems of increasing complexity. Two methods were developed: transformational analogy, with solutions to related problems are incremently transformed into the solution of a new problem, and derivational analogy with the problem solving strategies, rather than the resultant solutions, are transferred across like problems.
Jaime G. Carbonell
ISMIS1
1984 Is There Natural Language after Data Bases?
abstract
No abstract available.
Jaime G. Carbonell
COLING1
1984 Coping with Extragrammaticality
abstract
Practical natural language interfaces must exhibit robust behaviour in the presence of extragrammatical user input. This paper classifies different types of grammatical deviations and related phenomena at the lexical and sentential levels, discussing recovery strategies tailored to specific phenomena in the classification. Such strategies constitute a tool chest of computationally tractable methods for coping with extragrammaticality in restricted domain natural language. Some of the strategies have been tested and proven viable in existing parsers.
Jaime G. Carbonell, Philip J. Hayes
COLING1
1984 Approaches to machine learning
abstract
Abstract The field of machine learning strives to develop methods and techniques to automate the acquisition of new information, new skills, and new ways of organizing existing information. This article reviews the major approaches to machine learning in symbolic domains, illustrated with occasional paradigmatic examples.
Pat Langley, Jaime G. Carbonell
J. Am. Soc. Inf. Sci.2
1983 Derivational Analogy and Its Role in Problem Solving
Jaime G. Carbonell
AAAI1
1983 Discourse Pragmatics and Ellipsis Resolution in Task-Oriented Natural Language Interfaces
abstract
This paper reviews discourse phenomena that occur frequently in task-oriented man-machine dialogs, reporting on an empirical study that demonstrates the necessity of handling ellipsis, anaphora, extragrammaticality, inter-sentential metalanguage, and other abbreviatory devices in order to achieve convivial user interaction. Invariably, users prefer to generate terse or fragmentary utterances instead of longer, more complete "standalone" expressions, even when given clear instructions to the contrary. The XCALIBUR expert system interface is designed to meet these needs, including generalized ellipsis resolution by means of a rule-based caseframe method superior to previous semantic grammar approaches.
Jaime G. Carbonell
ACL1
1983 The XCALIBUR Project: A Natural Language Interface to Expert Systems
Jaime G. Carbonell, W. Mark Boggs, Michael L. Mauldin, Peter G. Anick
IJCAI1
1983 A Framework for Processing Corrections in Task-Oriented Dialogues
Philip J. Hayes, Jaime G. Carbonell
IJCAI2
1983 Recovery Strategies for Parsing Extragrammatical Language
Jaime G. Carbonell, Philip J. Hayes
Am. J. Comput. Linguistics1
1982 Experiential Learning in Analogical Problem Solving
Jaime G. Carbonell
AAAI1
1981 Dynamic Strategy Selection in Flexible Parsing
abstract
Robust natural language interpretation requires strong semantic domain models, "fail-soft" recovery heuristics, and very flexible control structures. Although single-strategy parsers have met with a measure of success, a multi-strategy approach is shown to provide a much higher degree of flexibility, redundancy, and ability to bring task-specific domain knowledge (in addition to general linguistic knowledge) to bear on both grammatical and ungrammatical input. A parsing algorithm is presented that integrates several different parsing strategies, with case-frame instantiation dominating. Each of these parsing strategies exploits different types of knowledge; and their combination provides a strong framework in which to process conjunctions, fragmentary input, and ungrammatical structures, as well as less exotic, grammatically correct input. Several specific heuristics for handling ungrammatical input are presented within this multi-strategy framework.
Jaime G. Carbonell, Philip J. Hayes
ACL1
1981 A Computational Model of Analogical Problem Solving
Jaime G. Carbonell
IJCAI1
1981 Multi-Strategy Construction-Specific Parsing for Flexible Data Base Query and Update
Philip J. Hayes, Jaime G. Carbonell
IJCAI2
1981 Counterplanning: A Strategy-Based Model of Adversary Planning in Real-World Situations
Jaime G. Carbonell
Artif. Intell.1
1981 Steps Toward Knowledge-Based Machine Translation
abstract
This paper considers the possibilities for knowledge-based automatic text translation in the light of recent advances in artificial intelligence. It is argued that competent translation requires some reasonable depth of understanding of the source text, and, in particular, access to detailed contextual information. The following machine translation paradigm is proposed. First, the source text is analyzed and mapped into a language-free conceptual representation. Inference mechanisms then apply contextual world knowledge to augment the representation in various ways, adding information about items that were only implicit in the input text. Finally, a natural-language generator maps appropriate sections of the language-free representation into the target language. We discuss several difficult translation problems from this viewpoint with examples of English-to-Spanish and English-to-Russian translations; and illustrate possible solutions as embodied in a computer understander called SAM, which reads certain kinds of newspaper stories, then summarizes or paraphrases them in a variety of languages.
Jaime G. Carbonell, Richard E. Cullingford, Anatole Gershman
IEEE Trans. Pattern Anal. Mach. Intell.1
1980 DELTA-MIN: A Search-Control Method for Information-Gathering Problems
Jaime G. Carbonell
AAAI1
1980 Metapher - A Key to Extensible Semantic Analysis
abstract
Interpreting metaphors is an integral and inescapable process in human understanding of natural language. This paper discusses a method of analyzing metaphors based on the existence of a small number of generalized metaphor mappings. Each generalized metaphor contains a recognition network, a basic mapping, additional transfer mappings, and an implicit intention component. It is argued that the method reduces metaphor interpretation from a reconstruction to a recognition task. Implications towards automating certain aspects of language learning are also discussed.
Jaime G. Carbonell
ACL1
1980 Towards a Process Model of Human Personality Traits
Jaime G. Carbonell
Artif. Intell.1
1979 Towards a Self-Extending Parser
abstract
This paper discusses an approach to incremental learning in natural language processing.The technique of projecting and integrating semantic constraints to learn word definitions is analyzed as Implemented in the POLITICS system.Extensions and improvements of this technique are developed.The problem of generalizing existing word meanings and understanding metaphorical uses of words Is addressed In terms of semantic constraint Integration.
Jaime G. Carbonell
ACL1
1979 Computer Models of Human Personality Traits
Jaime G. Carbonell
IJCAI1
1979 The Counterplanning Process: Reasoning under Adversity
Jaime G. Carbonell
IJCAI1
1978 Comments on the paper of Cherniavsky: "On artificial intelligence and attempts to disprove its existance"
Jaime G. Carbonell, Roger C. Schank
Inf. Syst.1