Prem Melville

dblp:71/2632 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
0since 2021 · last 2013
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 7 first-authorDatabases, data management, data science and information retrieval · 14 · 6 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Information extraction and text analysis · 51% Efficient and distributed learning · 19% Reinforcement learning · 14%
Databases, data mining, and information retrieval
7 papers
Data mining · 68% Data stream processing · 17% Machine learning and data management · 12%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational finance and economics · 100%

Topics — the 23 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
sentiment analysis
0.222009
Sentiment analysis of blogs by combining lexical knowledge with text classification · KDD 2009
Document-Word Co-regularization for Semi-supervised Sentiment Analysis · ICDM 2008
Data mining › representation learning
dictionary learning
0.112012
Online L1-Dictionary Learning with Application to Novel Document Detection · NIPS 2012
Machine learning › Efficient and distributed learning
active learning
0.122009
Uncertainty sampling and transductive experimental design for active dual supervision · ICML 2009
Diverse ensembles for active learning · ICML 2004
Knowledge, reasoning and agents › Knowledge representation and reasoning
ontology
0.112011
Concept Labeling: Building Text Classifiers with Minimal Supervision · IJCAI 2011
Natural language and speech › Information extraction and text analysis
text classification
0.112011
Concept Labeling: Building Text Classifiers with Minimal Supervision · IJCAI 2011
Natural language and speech › Information extraction and text analysis › text classification
weakly supervised text classification
0.112011
Concept Labeling: Building Text Classifiers with Minimal Supervision · IJCAI 2011
Machine learning › Reinforcement learning › markov decision process
constrained markov decision process
0.112010
Optimizing debt collections using constrained reinforcement learning · KDD 2010
Machine learning › Reinforcement learning
constrained reinforcement learning
0.112010
Optimizing debt collections using constrained reinforcement learning · KDD 2010
Machine learning and data management
active learning
0.122005
An Expected Utility Approach to Active Feature-Value Acquisition · ICDM 2005
Active Feature-Value Acquisition for Classifier Induction · ICDM 2004
Data mining › predictive modeling
classification
0.122005
An Expected Utility Approach to Active Feature-Value Acquisition · ICDM 2005
Active Feature-Value Acquisition for Classifier Induction · ICDM 2004
Data mining › feature engineering
feature-value acquisition
0.122005
An Expected Utility Approach to Active Feature-Value Acquisition · ICDM 2005
Active Feature-Value Acquisition for Classifier Induction · ICDM 2004
Machine learning › Efficient and distributed learning › active learning
uncertainty sampling
0.112009
Uncertainty sampling and transductive experimental design for active dual supervision · ICML 2009
Natural language and speech › Information extraction and text analysis › sentiment analysis
lexicon-based sentiment analysis
0.112008
Document-Word Co-regularization for Semi-supervised Sentiment Analysis · ICDM 2008
Natural language and speech › Information extraction and text analysis › sentiment analysis › sentiment classification
semi-supervised sentiment classification
0.112008
Document-Word Co-regularization for Semi-supervised Sentiment Analysis · ICDM 2008
Data mining › business intelligence
customer targeting
0.112008
Customer targeting models using actively-selected web content · KDD 2008
Machine learning › Kernel, tree and ensemble methods › ensemble learning
ensemble diversity
0.122004
Constructing Diverse Classifier Ensembles using Artificial Training Examples · IJCAI 2003
Diverse ensembles for active learning · ICML 2004
Machine learning › Kernel, tree and ensemble methods
ensemble learning
0.122004
Constructing Diverse Classifier Ensembles using Artificial Training Examples · IJCAI 2003
Diverse ensembles for active learning · ICML 2004
Data mining › predictive modeling › classification
cost-sensitive learning
0.112005
An Expected Utility Approach to Active Feature-Value Acquisition · ICDM 2005
Machine learning › Efficient and distributed learning › active learning › disagreement-based active learning
query by committee
0.012004
Diverse ensembles for active learning · ICML 2004
Data integration and cleaning
missing data
0.022005
An Expected Utility Approach to Active Feature-Value Acquisition · ICDM 2005
Active Feature-Value Acquisition for Classifier Induction · ICDM 2004
Data mining › text mining
text classification
0.012009
Uncertainty sampling and transductive experimental design for active dual supervision · ICML 2009
Data mining
semi-supervised learning
0.012008
Document-Word Co-regularization for Semi-supervised Sentiment Analysis · ICDM 2008
Machine learning › Deep learning architectures and training
data augmentation
0.012003
Constructing Diverse Classifier Ensembles using Artificial Training Examples · IJCAI 2003

Methods — techniques the papers use, named apart from their topics

active learning · 0.3domain knowledge integration · 0.3data mining · 0.3reinforcement learning · 0.2constrained markov decision process · 0.2graph-based learning · 0.2experimental design · 0.2sublinear regret analysis · 0.1alternating direction method of multipliers · 0.1transfer learning · 0.1semi-supervised learning · 0.1multi-task learning · 0.1uncertainty sampling · 0.1supervised learning · 0.1simulation · 0.1lexical prior knowledge · 0.1bipartite graph co-regularization · 0.1
YearPublicationVenuePosition
2013 Amplifying the voice of youth in Africa via text analytics
abstract
U-report is an open-source SMS platform operated by UNICEF Uganda, designed to give community members a voice on issues that impact them. Data received by the system are either SMS responses to a poll conducted by UNICEF, or unsolicited reports of a problem occurring within the community. There are currently 200,000 U-report participants, and they send up to 10,000 unsolicited text messages a week. The objective of the program in Uganda is to understand the data in real-time, and have issues addressed by the appropriate department in UNICEF in a timely manner. Given the high volume and velocity of the data streams, manual inspection of all messages is no longer sustainable. This paper describes an automated message-understanding and routing system deployed by IBM at UNICEF. We employ recent advances in data mining to get the most out of labeled training data, while incorporating domain knowledge from experts. We discuss the trade-offs, design choices and challenges in applying such techniques in a real-world deployment.
Prem Melville, Vijil Chenthamarakshan, Richard D. Lawrence, James Powell, Moses Mugisha, Sharad Sapra, Rajesh Anandan, Solomon Assefa
KDD1
2012 Learning to rank for robust question answering
abstract
This paper aims to solve the problem of improving the ranking of answer candidates for factoid based questions in a state-of-the-art Question Answering system. We first provide an extensive comparison of 5 ranking algorithms on two datasets -- from the Jeopardy quiz show and a medical domain. We then show the effectiveness of a cascading approach, where the ranking produced by one ranker is used as input to the next stage. The cascading approach shows sizeable gains on both datasets. We finally evaluate several rank aggregation techniques to combine these algorithms, and find that Supervised Kemeny aggregation is a robust technique that always beats the baseline ranking approach used by Watson for the Jeopardy competition. We further corroborate our results on TREC Question Answering datasets.
Arvind Agarwal, Hema Raghavan, Karthik Subbian, Prem Melville, Richard D. Lawrence, David Gondek, James Fan
CIKM4
2012 Online L1-Dictionary Learning with Application to Novel Document Detection
abstract
Given their pervasive use, social media, such as Twitter, have become a leading source of breaking news. A key task in the automated identification of such news is the detection of novel documents from a voluminous stream of text documents in a scalable manner. Motivated by this challenge, we introduce the problem of online L1-dictionary learning where unlike traditional dictionary learning, which uses squared loss, the L1-penalty is used for measuring the reconstruction error. We present an efficient online algorithm for this problem based on alternating directions method of multipliers, and establish a sublinear regret bound for this algorithm. Empirical results on news-stream and Twitter data, shows that this online L1-dictionary learning algorithm for novel document detection gives more than an order of magnitude speedup over the previously known batch algorithm, without any significant loss in quality of results. Our algorithm for online L1-dictionary learning could be of independent interest.
Shiva Prasad Kasiviswanathan, Huahua Wang, Arindam Banerjee 0001, Prem Melville
NIPS4
2011 Emerging topic detection using dictionary learning
abstract
Streaming user-generated content in the form of blogs, microblogs, forums, and multimedia sharing sites, provides a rich source of data from which invaluable information and insights maybe gleaned. Given the vast volume of such social media data being continually generated, one of the challenges is to automatically tease apart the emerging topics of discussion from the constant background chatter. Such emerging topics can be identified by the appearance of multiple posts on a unique subject matter, which is distinct from previous online discourse. We address the problem of identifying emerging topics through the use of dictionary learning. We propose a two stage approach respectively based on detection and clustering of novel user-generated content. We derive a scalable approach by using the alternating directions method to solve the resulting optimization problems. Empirical results show that our proposed approach is more effective than several baselines in detecting emerging topics in traditional news story and newsgroup data. We also demonstrate the practical application to social media analysis, based on a study on streaming data from Twitter.
Shiva Prasad Kasiviswanathan, Prem Melville, Arindam Banerjee 0001, Vikas Sindhwani
CIKM2
2011 Concept Labeling: Building Text Classifiers with Minimal Supervision
abstract
The rapid construction of supervised text classification models is becoming a pervasive need across many modern applications. To reduce human-labeling bottlenecks, many new statistical paradigms (e.g., active, semi-supervised, transfer and multi-task learning) have been vigorously pursued in recent literature with varying degrees of empirical success. Concurrently, the emergence of Web 2.0 platforms in the last decade has enabled a world-wide, collaborative human effort to construct a massive ontology of concepts with very rich, detailed and accurate descriptions. In this paper we propose a new framework to extract supervisory information from such ontologies and complement it with a shift in human effort from direct labeling of examples in the domain of interest to the much more efficient identification of concept-class associations. Through empirical studies on text categorization problems using the Wikipedia ontology, we show that this shift allows very high-quality models to be immediately induced at virtually no cost. 1
Vijil Chenthamarakshan, Prem Melville, Vikas Sindhwani, Richard D. Lawrence
IJCAI2
2010 Optimizing debt collections using constrained reinforcement learning
abstract
The problem of optimally managing the collections process by taxation authorities is one of prime importance, not only for the revenue it brings but also as a means to administer a fair taxing system. The analogous problem of debt collections management in the private sector, such as banks and credit card companies, is also increasingly gaining attention. With the recent successes in the applications of data analytics and optimization to various business areas, the question arises to what extent such collections processes can be improved by use of leading edge data modeling and optimization techniques. In this paper, we propose and develop a novel approach to this problem based on the framework of constrained Markov Decision Process (MDP), and report on our experience in an actual deployment of a tax collections optimization system at New York State Department of Taxation and Finance (NYS DTF).
Naoki Abe, Prem Melville, Cezar Pendus, Chandan K. Reddy, David L. Jensen, Vince P. Thomas, James J. Bennett, Gary F. Anderson, Brent R. Cooley, Melissa Kowalczyk, Mark Domick, Timothy Gardinier
KDD2
2010 A Unified Approach to Active Dual Supervision for Labeling Features and Examples
Josh Attenberg, Prem Melville, Foster J. Provost
ECML/PKDD (1)2
2010 Medical data mining: insights from winning two competitions
Saharon Rosset, Claudia Perlich, Grzegorz Swirszcz, Prem Melville, Yan Liu 0002
Data Min. Knowl. Discov.4
2009 Uncertainty sampling and transductive experimental design for active dual supervision
abstract
Dual supervision refers to the general setting of learning from both labeled examples as well as labeled features. Labeled features are naturally available in tasks such as text classification where it is frequently possible to provide domain knowledge in the form of words that associate strongly with a class. In this paper, we consider the novel problem of active dual supervision, or, how to optimally query an example and feature labeling oracle to simultaneously collect two different forms of supervision, with the objective of building the best classifier in the most cost effective manner. We apply classical uncertainty and experimental design based active learning schemes to graph/kernel-based dual supervision models. Empirical studies confirm the potential of these schemes to significantly reduce the cost of acquiring labeled data for training high-quality models.
Vikas Sindhwani, Prem Melville, Richard D. Lawrence
ICML2
2009 Sentiment analysis of blogs by combining lexical knowledge with text classification
abstract
The explosion of user-generated content on the Web has led to new opportunities and significant challenges for companies, that are increasingly concerned about monitoring the discussion around their products. Tracking such discussion on weblogs, provides useful insight on how to improve products or market them more effectively. An important component of such analysis is to characterize the sentiment expressed in blogs about specific brands and products. Sentiment Analysis focuses on this task of automatically identifying whether a piece of text expresses a positive or negative opinion about the subject matter. Most previous work in this area uses prior lexical knowledge in terms of the sentiment-polarity of words. In contrast, some recent approaches treat the task as a text classification problem, where they learn to classify sentiment based only on labeled training data. In this paper, we present a unified framework in which one can use background lexical information in terms of word-class associations, and refine this information for specific domains using any available training examples. Empirical results on diverse domains show that our approach performs better than using background knowledge or training data in isolation, as well as alternative approaches to using lexical knowledge with text classification.
Prem Melville, Wojciech Gryc, Richard D. Lawrence
KDD1
2008 Document-Word Co-regularization for Semi-supervised Sentiment Analysis
abstract
The goal of sentiment prediction is to automatically identify whether a given piece of text expresses positive or negative opinion towards a topic of interest. One can pose sentiment prediction as a standard text categorization problem, but gathering labeled data turns out to be a bottleneck. Fortunately, background knowledge is often available in the form of prior information about the sentiment polarity of words in a lexicon. Moreover, in many applications abundant unlabeled data is also available. In this paper, we propose a novel semi-supervised sentiment prediction algorithm that utilizes lexical prior knowledge in conjunction with unlabeled examples. Our method is based on joint sentiment analysis of documents and words based on a bipartite graph representation of the data. We present an empirical study on a diverse collection of sentiment prediction problems which confirms that our semi-supervised lexical models significantly outperform purely supervised and competing semi-supervised techniques.
Vikas Sindhwani, Prem Melville
ICDM2
2008 Customer targeting models using actively-selected web content
abstract
We consider the problem of predicting the likelihood that a company will purchase a new product from a seller. The statistical models we have developed at IBM for this purpose rely on historical transaction data coupled with structured firmographic information like the company revenue, number of employees and so on. In this paper, we extend this methodology to include additional text-based features based on analysis of the content on each company's website. Empirical results demonstrate that incorporating such web content can significantly improve customer targeting. Furthermore, we present methods to actively select only the web content that is likely to improve our models, while reducing the costs of acquisition and processing.
Prem Melville, Saharon Rosset, Richard D. Lawrence
KDD1
2008 Using predictive analysis to improve invoice-to-cash collection
abstract
It is commonly agreed that accounts receivable (AR) can be a source of financial difficulty for firms when they are not efficiently managed and are underperforming. Experience across multiple industries shows that effective management of AR and overall financial performance of firms are positively correlated. In this paper we address the problem of reducing outstanding receivables through improvements in the collections strategy. Specifically, we demonstrate how supervised learning can be used to build models for predicting the payment outcomes of newly-created invoices, thus enabling customized collection actions tailored for each invoice or customer. Our models can predict with high accuracy if an invoice will be paid on time or not and can provide estimates of the magnitude of the delay. We illustrate our techniques in the context of real-world transaction data from multiple firms. Finally, simulation results show that our approach can reduce collection time up to a factor of four compared to a baseline that is not model-driven.
Sai Zeng, Prem Melville, Christian A. Lang, Ioana M. Boier-Martin, Conrad Murphy
KDD2
2007 Data acquisition and cost-effective predictive modeling: targeting offers for electronic commerce
abstract
Electronic commerce is revolutionizing the way we think about data modeling, by making it possible to integrate the processes of (costly) data acquisition and model induction. The opportunity for improving modeling through costly data acquisition presents itself for a diverse set of electronic commerce modeling tasks, from personalization to customer lifetime value modeling; we illustrate with the running example of choosing offers to display to web-site visitors, which captures important aspects in a familiar setting. Considering data acquisition costs explicitly can allow the building of predictive models at significantly lower costs, and a modeler may be able to improve performance via new sources of information that previously were too expensive to consider. However, existing techniques for integrating modeling and data acquisition cannot deal with the rich environment that electronic commerce presents. We discuss several possible data acquisition settings, the challenges involved in the integration with modeling, and various research areas that may supply parts of an ultimate solution. We also present and demonstrate briefly a unified framework within which one can integrate acquisitions of different types, with any cost structure and any predictive modeling objective.
Foster J. Provost, Prem Melville, Maytal Saar-Tsechansky
ICEC2
2005 Active Learning for Probability Estimation Using Jensen-Shannon Divergence
Prem Melville, Stewart M. Yang, Maytal Saar-Tsechansky, Raymond J. Mooney
ECML1
2005 Combining Bias and Variance Reduction Techniques for Regression Trees
Yuk Lai Suen, Prem Melville, Raymond J. Mooney
ECML2
2005 An Expected Utility Approach to Active Feature-Value Acquisition
abstract
In many classification tasks, training data have missing feature values that can be acquired at a cost. For building accurate predictive models, acquiring all missing values is often prohibitively expensive or unnecessary, while acquiring a random subset of feature values may not be most effective. The goal of active feature-value acquisition is to incrementally select feature values that are most cost-effective for improving the model's accuracy. We present an approach that acquires feature values for inducing a classification model based on an estimation of the expected improvement in model accuracy per unit cost. Experimental results demonstrate that our approach consistently reduces the cost of producing a model of a desired accuracy compared to random feature acquisitions.
Prem Melville, Foster J. Provost, Raymond J. Mooney
ICDM1
2004 Active Feature-Value Acquisition for Classifier Induction
abstract
Many induction problems include missing data that can be acquired at a cost. For building accurate predictive models, acquiring complete information for all instances is often expensive or unnecessary, while acquiring information for a random subset of instances may not be most effective. Active feature-value acquisition tries to reduce the cost of achieving a desired model accuracy by identifying instances for which obtaining complete information is most informative. We present an approach in which instances are selected for acquisition based on the current model's accuracy and its confidence in the prediction. Experimental results demonstrate that our approach can induce accurate models using substantially fewer feature-value acquisitions as compared to alternative policies.
Prem Melville, Maytal Saar-Tsechansky, Foster J. Provost, Raymond J. Mooney
ICDM1
2004 Diverse ensembles for active learning
abstract
Query by Committee is an effective approach to selective sampling in which disagreement amongst an ensemble of hypotheses is used to select data for labeling. Query by Bagging and Query by Boosting are two practical implementations of this approach that use Bagging and Boosting, respectively, to build the committees. For effective active learning, it is critical that the committee be made up of consistent hypotheses that are very different from each other. DECORATE is a recently developed method that directly constructs such diverse committees using artificial training data. This paper introduces ACTIVE-DECORATE, which uses DECORATE committees to select good training examples. Extensive experimental results demonstrate that, in general, ACTIVE-DECORATE outperforms both Query by Bagging and Query by Boosting.
Prem Melville, Raymond J. Mooney
ICML1
2003 Constructing Diverse Classifier Ensembles using Artificial Training Examples
Prem Melville, Raymond J. Mooney
IJCAI1