Richard D. Lawrence

dblp:33/2948 · DBLP profile ↗
← Back
22ranked-venue papers
3as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 1 first-authorDatabases, data management, data science and information retrieval · 15 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Information extraction and text analysis · 30% Learning paradigms · 19% Transfer learning and domain adaptation · 14%
Databases, data mining, and information retrieval
8 papers
Data mining · 81% Machine learning and data management · 19%

Topics — the 19 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning paradigms
multiple instance learning
0.322013
MI2LS: multi-instance learning from multiple informationsources · KDD 2013
MILEAGE: Multiple Instance LEArning with Global Embedding · ICML (3) 2013
Machine learning › Learning theory
generalization bounds
0.212013
MILEAGE: Multiple Instance LEArning with Global Embedding · ICML (3) 2013
Machine learning › Representation and self-supervised learning
multi-view learning
0.212013
MI2LS: multi-instance learning from multiple informationsources · KDD 2013
Data mining › text mining
text classification
0.222011
Transfer Latent Semantic Learning: Microblog Mining with Less Supervision · AAAI 2011
Uncertainty sampling and transductive experimental design for active dual supervision · ICML 2009
Machine learning › Transfer learning and domain adaptation
multi-view transfer learning
0.112011
Multi-view transfer learning with a large margin approach · KDD 2011
Knowledge, reasoning and agents › Knowledge representation and reasoning
ontology
0.112011
Concept Labeling: Building Text Classifiers with Minimal Supervision · IJCAI 2011
Natural language and speech › Information extraction and text analysis
text classification
0.112011
Concept Labeling: Building Text Classifiers with Minimal Supervision · IJCAI 2011
Natural language and speech › Information extraction and text analysis › text classification
weakly supervised text classification
0.112011
Concept Labeling: Building Text Classifiers with Minimal Supervision · IJCAI 2011
Machine learning and data management › weak supervision
multiple instance learning
0.112011
Multiple Instance Learning on Structured Data · NIPS 2011
Machine learning › Efficient and distributed learning
active learning
0.112009
Uncertainty sampling and transductive experimental design for active dual supervision · ICML 2009
Natural language and speech › Information extraction and text analysis
sentiment analysis
0.112009
Sentiment analysis of blogs by combining lexical knowledge with text classification · KDD 2009
Machine learning › Efficient and distributed learning › active learning
uncertainty sampling
0.112009
Uncertainty sampling and transductive experimental design for active dual supervision · ICML 2009
Data mining › business intelligence
customer targeting
0.112008
Customer targeting models using actively-selected web content · KDD 2008
Data mining › predictive modeling › classification › class imbalance
rare category detection
0.112008
Graph-Based Rare Category Detection · ICDM 2008
Data mining
anomaly detection
0.022008
Graph-Based Rare Category Detection · ICDM 2008
High-quantile modeling for customer wallet estimation and other applications · KDD 2007
Data mining
predictive modeling
0.012003
Passenger-based predictive modeling of airline no-show rates · KDD 2003
Mathematical optimization
nonconvex optimization
0.012011
Multiple Instance Learning on Structured Data · NIPS 2011
Data mining › anomaly detection
fraud detection
0.012007
High-quantile modeling for customer wallet estimation and other applications · KDD 2007
Computational finance and economics
revenue management
0.012003
Passenger-based predictive modeling of airline no-show rates · KDD 2003

Methods — techniques the papers use, named apart from their topics

domain knowledge integration · 0.3data mining · 0.3latent semantic learning · 0.2hinge loss · 0.2concave-convex constraint programming · 0.2active learning · 0.2non-convex optimization · 0.2multi-view feature fusion · 0.2large margin method · 0.2consistency regularization · 0.2bundle method · 0.2cutting-plane method · 0.1cutting plane method · 0.1uncertainty sampling · 0.1graph-based learning · 0.1experimental design · 0.1graph similarity · 0.1naive bayes · 0.0
YearPublicationVenuePosition
2015 A general framework for predictive tensor modeling with domain knowledge
Yada Zhu, Jingrui He, Richard D. Lawrence
Data Min. Knowl. Discov.3
2013 MILEAGE: Multiple Instance LEArning with Global Embedding
abstract
Multiple Instance Learning (MIL) methods generally represent each example as a collection of instances such that the features for local objects can be better captured, whereas traditional learning methods typically extract a global feature vector for each example as an integral part. However, there is limited research work on which of the two learning scenarios performs better. This paper proposes a novel framework – \emphMultiple Instance LEArning with Global Embedding (MILEAGE), in which the global feature vectors for traditional learning methods are integrated into the MIL setting. MILEAGE can leverage the benefits derived from both learning settings. Within the proposed framework, a large margin method is formulated. In particular, the proposed method adaptively tunes the weights on the two different kinds of feature representations (i.e., global and multiple instance) for each example and trains the classifier simultaneously. An alternative algorithm is proposed to solve the resulting optimization problem, which extends the bundle method to the non-convex case. Some important properties of the proposed method, such as the convergence rate and the generalization error rate, are analyzed. A series of experiments have been conducted to demonstrate the advantages of the proposed method over several state-of-the-art multiple instance and traditional learning methods.
Dan Zhang 0007, Jingrui He, Luo Si, Richard D. Lawrence
ICML (3)4
2013 Amplifying the voice of youth in Africa via text analytics
abstract
U-report is an open-source SMS platform operated by UNICEF Uganda, designed to give community members a voice on issues that impact them. Data received by the system are either SMS responses to a poll conducted by UNICEF, or unsolicited reports of a problem occurring within the community. There are currently 200,000 U-report participants, and they send up to 10,000 unsolicited text messages a week. The objective of the program in Uganda is to understand the data in real-time, and have issues addressed by the appropriate department in UNICEF in a timely manner. Given the high volume and velocity of the data streams, manual inspection of all messages is no longer sustainable. This paper describes an automated message-understanding and routing system deployed by IBM at UNICEF. We employ recent advances in data mining to get the most out of labeled training data, while incorporating domain knowledge from experts. We discuss the trade-offs, design choices and challenges in applying such techniques in a real-world deployment.
Prem Melville, Vijil Chenthamarakshan, Richard D. Lawrence, James Powell, Moses Mugisha, Sharad Sapra, Rajesh Anandan, Solomon Assefa
KDD3
2013 MI2LS: multi-instance learning from multiple informationsources
abstract
In Multiple Instance Learning (MIL), each entity is normally expressed as a set of instances. Most of the current MIL methods only deal with the case when each instance is represented by one type of features. However, in many real world applications, entities are often described from several different information sources/views. For example, when applying MIL to image categorization, the characteristics of each image can be derived from both its RGB features and SIFT features. Previous research work has shown that, in traditional learning methods, leveraging the consistencies between different information sources could improve the classification performance drastically.
Dan Zhang 0007, Jingrui He, Richard D. Lawrence
KDD3
2012 Learning to rank for robust question answering
abstract
This paper aims to solve the problem of improving the ranking of answer candidates for factoid based questions in a state-of-the-art Question Answering system. We first provide an extensive comparison of 5 ranking algorithms on two datasets -- from the Jeopardy quiz show and a medical domain. We then show the effectiveness of a cascading approach, where the ranking produced by one ranker is used as input to the next stage. The cascading approach shows sizeable gains on both datasets. We finally evaluate several rank aggregation techniques to combine these algorithms, and find that Supervised Kemeny aggregation is a robust technique that always beats the baseline ranking approach used by Watson for the Jeopardy competition. We further corroborate our results on TREC Question Answering datasets.
Arvind Agarwal, Hema Raghavan, Karthik Subbian, Prem Melville, Richard D. Lawrence, David Gondek, James Fan
CIKM5
2011 Transfer Latent Semantic Learning: Microblog Mining with Less Supervision
abstract
The increasing volume of information generated on micro-blogging sites such as Twitter raises several challenges to traditional text mining techniques. First, most texts from those sites are abbreviated due to the constraints of limited characters in one post; second, the input usually comes in streams of large-volumes. Therefore, it is of significant importance to develop effective and efficient representations of abbreviated texts for better filtering and mining. In this paper, we introduce a novel transfer learning approach, namely transfer latent semantic learning, that utilizes a large number of related tagged documents with rich information from other sources (source domain) to help build a robust latent semantic space for the abbreviated texts (target domain). This is achieved by simultaneously minimizing the document reconstruction error and the classification error of the labeled examples from the source domain by building a classifier with hinge loss in the latent semantic space. We demonstrate the effectiveness of our method by applying them to the task of classifying and tagging abbreviated texts. Experimental results on both synthetic datasets and real application datasets, including Reuters-21578 and Twitter data, suggest substantial improvements using our approach over existing ones.
Dan Zhang 0007, Yan Liu 0002, Richard D. Lawrence, Vijil Chenthamarakshan
AAAI3
2011 Concept Labeling: Building Text Classifiers with Minimal Supervision
abstract
The rapid construction of supervised text classification models is becoming a pervasive need across many modern applications. To reduce human-labeling bottlenecks, many new statistical paradigms (e.g., active, semi-supervised, transfer and multi-task learning) have been vigorously pursued in recent literature with varying degrees of empirical success. Concurrently, the emergence of Web 2.0 platforms in the last decade has enabled a world-wide, collaborative human effort to construct a massive ontology of concepts with very rich, detailed and accurate descriptions. In this paper we propose a new framework to extract supervisory information from such ontologies and complement it with a shift in human effort from direct labeling of examples in the domain of interest to the much more efficient identification of concept-class associations. Through empirical studies on text categorization problems using the Wikipedia ontology, we show that this shift allows very high-quality models to be immediately induced at virtually no cost. 1
Vijil Chenthamarakshan, Prem Melville, Vikas Sindhwani, Richard D. Lawrence
IJCAI4
2011 Multi-view transfer learning with a large margin approach
abstract
Transfer learning has been proposed to address the problem of scarcity of labeled data in the target domain by leveraging the data from the source domain. In many real world applications, data is often represented from different perspectives, which correspond to multiple views. For example, a web page can be described by its contents and its associated links. However, most existing transfer learning methods fail to capture the multi-view {nature}, and might not be best suited for such applications.
Dan Zhang 0007, Jingrui He, Yan Liu 0002, Luo Si, Richard D. Lawrence
KDD5
2011 Multiple Instance Learning on Structured Data
abstract
Most existing Multiple-Instance Learning (MIL) algorithms assume data instances and/or data bags are independently and identically distributed. But there often exists rich additional dependency/structure information between instances/bags within many applications of MIL. Ignoring this structure information limits the performance of existing MIL algorithms. This paper explores the research problem as multiple instance learning on structured data (MILSD) and formulates a novel framework that considers additional structure information. In particular, an effective and efficient optimization algorithm has been proposed to solve the original non-convex optimization problem by using a combination of Concave-Convex Constraint Programming (CCCP) method and an adapted Cutting Plane method, which deals with two sets of constraints caused by learning on instances within individual bags and learning on structured data. Our method has the nice convergence property, with specified precision on each set of constraints. Experimental results on three different applications, i.e., webpage classification, market targeting, and protein fold identification, clearly demonstrate the advantages of the proposed method over state-of-the-art methods.
Dan Zhang 0007, Yan Liu 0002, Luo Si, Jian Zhang 0003, Richard D. Lawrence
NIPS5
2009 Graph-based transfer learning
abstract
Transfer learning is the task of leveraging the information from labeled examples in some domains to predict the labels for examples in another domain. It finds abundant practical applications, such as sentiment prediction, image classification and network intrusion detection. In this paper, we propose a graph-based transfer learning framework. It propagates the label information from the source domain to the target domain via the example-feature-example tripartite graph, and puts more emphasis on the labeled examples from the target domain via the example-example bi-partite graph. Our framework is semi-supervised and non-parametric in nature and thus more flexible. We also develop an iterative algorithm so that our framework is scalable to large-scale applications. It enjoys the theoretical property of convergence. Compared with existing transfer learning methods, the proposed framework propagates the label information to both the features irrelevant to the source domain and the unlabeled examples in the target omain via the common features in a principled way. Experimental results on 3 real data sets demonstrate the effectiveness of our algorithm.
Jingrui He, Yan Liu 0002, Richard D. Lawrence
CIKM3
2009 Uncertainty sampling and transductive experimental design for active dual supervision
abstract
Dual supervision refers to the general setting of learning from both labeled examples as well as labeled features. Labeled features are naturally available in tasks such as text classification where it is frequently possible to provide domain knowledge in the form of words that associate strongly with a class. In this paper, we consider the novel problem of active dual supervision, or, how to optimally query an example and feature labeling oracle to simultaneously collect two different forms of supervision, with the objective of building the best classifier in the most cost effective manner. We apply classical uncertainty and experimental design based active learning schemes to graph/kernel-based dual supervision models. Empirical studies confirm the potential of these schemes to significantly reduce the cost of acquiring labeled data for training high-quality models.
Vikas Sindhwani, Prem Melville, Richard D. Lawrence
ICML3
2009 Sentiment analysis of blogs by combining lexical knowledge with text classification
abstract
The explosion of user-generated content on the Web has led to new opportunities and significant challenges for companies, that are increasingly concerned about monitoring the discussion around their products. Tracking such discussion on weblogs, provides useful insight on how to improve products or market them more effectively. An important component of such analysis is to characterize the sentiment expressed in blogs about specific brands and products. Sentiment Analysis focuses on this task of automatically identifying whether a piece of text expresses a positive or negative opinion about the subject matter. Most previous work in this area uses prior lexical knowledge in terms of the sentiment-polarity of words. In contrast, some recent approaches treat the task as a text classification problem, where they learn to classify sentiment based only on labeled training data. In this paper, we present a unified framework in which one can use background lexical information in terms of word-class associations, and refine this information for specific domains using any available training examples. Empirical results on diverse domains show that our approach performs better than using background knowledge or training data in isolation, as well as alternative approaches to using lexical knowledge with text classification.
Prem Melville, Wojciech Gryc, Richard D. Lawrence
KDD3
2008 Graph-Based Rare Category Detection
abstract
Rare category detection is the task of identifying examples from rare classes in an unlabeled data set. It is an open challenge in machine learning and plays key roles in real applications such as financial fraud detection, network intrusion detection, astronomy, spam image detection, etc. In this paper, we develop a new graph-based method for rare category detection named GRADE. It makes use of the global similarity matrix motivated by the manifold ranking algorithm, which results in more compact clusters for the minority classes; by selecting examples from the regions where probability density changes the most, it relaxes the assumption that the majority classes and the minority classes are separable. Furthermore, when detailed information about the data set is not available, we develop a modified version of GRADE named GRADE-LI, which only needs an upper bound on the proportion of each minority class as input. Besides working with data with structured features, both GRADE and GRADE-LI can also work with graph data, which can not be handled by existing rare category detection methods. Experimental results on both synthetic and real data sets demonstrate the effectiveness of the GRADE and GRADE-LI algorithms.
Jingrui He, Yan Liu 0002, Richard D. Lawrence
ICDM3
2008 Customer targeting models using actively-selected web content
abstract
We consider the problem of predicting the likelihood that a company will purchase a new product from a seller. The statistical models we have developed at IBM for this purpose rely on historical transaction data coupled with structured firmographic information like the company revenue, number of employees and so on. In this paper, we extend this methodology to include additional text-based features based on analysis of the content on each company's website. Empirical results demonstrate that incorporating such web content can significantly improve customer targeting. Furthermore, we present methods to actively select only the web content that is likely to improve our models, while reducing the costs of acquisition and processing.
Prem Melville, Saharon Rosset, Richard D. Lawrence
KDD3
2007 High-quantile modeling for customer wallet estimation and other applications
abstract
In this paper we discuss the important practical problem of customer wallet estimation, i.e., estimation of potential spending by customers(rather than their expected spending). For this purpose we utilize quantile modeling, whose goal is to estimate a quantile of the discriminative conditional distribution of the response, rather than the mean, which is the implicit goal of most standard regression approaches. We argue that a notion of wallet can be captured through high quantile modeling (e.g, estimating the 90th percentile), and describe a wallet estimation implementation within IBM's Market Alignment Program (MAP). We also discuss the wide range of domains where high-quantile modeling can be practically important: estimating opportunities in sales and marketing domains, defining 'surprising' patterns for outlier and fraud detection and more. We survey some existing approaches for quantile modeling, and propose adaptations of nearest-neighbor and regression-tree approaches to quantile modeling. We demonstrate the various models' performance in high quantile estimation in several domains, including our motivating problem of estimating the 'realistic' IT wallets of IBM customers.
Claudia Perlich, Saharon Rosset, Richard D. Lawrence, Bianca Zadrozny
KDD3
2006 Data-Enhanced Predictive Modeling for Sales Targeting
abstract
We describe and analyze the idea of data-enhanced predictive modeling (DEM). The term “enhanced” here refers to the case that the data used for modeling is sampled not from the true target population, but from an alternative (closely related) population, from which much larger samples are available. This leads to a “bias-variance” tradeoff, which implies that in some cases, DEM can improve predictive performance on the true target population. We theoretically analyze this tradeoff for the case of linear regression. We illustrate DEM on a problem of sales targeting for a set of software products. The “correct” learning problem is to differentiate non-customers from newly acquired customers. The latter, however, are scarce. We illustrate how we can build better prediction models by using more flexible definitions of interesting targets, which give bigger learning samples.
Saharon Rosset, Richard D. Lawrence
SDM2
2003 Passenger-based predictive modeling of airline no-show rates
abstract
Airlines routinely overbook flights based on the expectation that some fraction of booked passengers will not show for each flight. Accurate forecasts of the expected number of no-shows for each flight can increase airline revenue by reducing the number of spoiled seats (empty seats that might otherwise have been sold) and the number of involuntary denied boardings at the departure gate. Conventional no-show forecasting methods typically average the no-show rates of historically similar flights, without the use of passenger-specific information.We develop two classes of models to predict cabin-level no-show rates using specific information on the individual passengers booked on each flight. The first of these models computes the no-show probability for each passenger, using both the cabin-level historical forecast and the extracted passenger features as explanatory variables. This passenger-level model is implemented using three different predictive methods: a C4.5 decision-tree, a segmented Naive Bayes algorithm, and a new aggregation method for an ensemble of probabilistic models. The second cabin-level model is formulated using the desired cabin-level no-show rate as the response variable. Inputs to this model include the predicted cabin-level no-show rates derived from the various passenger-level models, as well as simple statistics of the features of the cabin passenger population. The cabin-level model is implemented using either linear regression, or as a direct probability model with explicit incorporation of the cabin-level no-show rates derived from the passenger-level model outputs.The new passenger-based models are compared to a conventional historical model, using train and evaluation data sets taken from over 1 million passenger name records. Standard metrics such as lift curves and mean-square cabin-level errors establish the improved accuracy of the passenger-based models over the historical model. All models are also evaluated using a simple revenue model, and it is shown that the cabin-level passenger-based model can produce between 0.4% and 3.2% revenue gain over the conventional model, depending on the revenue-model parameters.
Richard D. Lawrence, Se June Hong, Jacques Cherrier
KDD1
2001 Personalization of Supermarket Product Recommendations
Richard D. Lawrence, George S. Almási, Vladimir Kotlyar, Marisa S. Viveros, Sastry Duri
Data Min. Knowl. Discov.1
1999 A Case Study in Information Delivery to Mass Retail Markets
Vladimir Kotlyar, Marisa S. Viveros, Sastry Duri, Richard D. Lawrence, George S. Almási
DEXA4
1999 A Scalable Parallel Algorithm for Self-Organizing Maps with Applications to Sparse Data Mining Problems
Richard D. Lawrence, George S. Almási, Holly E. Rushmeier
Data Min. Knowl. Discov.1
1997 Visualizing customer segmentations produce by self organizing maps (case study)
abstract
We describe a set of visualization programs developed for understanding segmentations of customer records produced by a self organizing map (SOM) algorithm. A SOM produces segments of similar customer records that can then be used as the basis of a marketing campaign. Since the characteristics that each segment will have in common are not specified a priori, visualization is essential to understanding the segment to design specific marketing strategies. Two different styles of visualizations were found to be useful for the two types of observers of the data. Abstract overviews of the entire segmentation were designed for analysts applying the SOM algorithm. Detailed scatterplots of individual records were designed for communicating the results to decision makers specifying marketing strategy.
Holly E. Rushmeier, Richard D. Lawrence, George S. Almási
IEEE Visualization2
1994 The IBM External User Interface for Scalable Parallel Systems
Vasanth Bala, Jehoshua Bruck, Raymond Bryant, Robert Cypher, Peter de Jong, Pablo Elustondo, Daniel D. Frye, Alex Ho, C. T. Howard Ho, Gail Irwin, Shlomo Kipnis, Richard D. Lawrence, Marc Snir
Parallel Comput.12