Sholom M. Weiss

dblp:w/SholomMWeiss · DBLP profile ↗
← Back
47ranked-venue papers
27as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 26 first-authorDatabases, data management, data science and information retrieval · 13 · 6 first-authorGraphics, computer vision, multimedia, augmented reality and games · 12 · 10 first-authorSystems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
9 papers
Data mining · 89% Information retrieval · 11%
Interdisciplinary, comprehensive, and emerging computing
7 papers
Computational science and engineering · 83% Bioinformatics and computational biology · 10% Computational finance and economics · 5%
Artificial intelligence
21 papers
Knowledge representation and reasoning · 58% Learning theory · 22% Kernel, tree and ensemble methods · 16%

Topics — the 30 heaviest of 41, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › text mining
text classification
0.132002
Experiments in high-dimensional text categorization · SIGIR 2002
A system for real-time competitive market intelligence · KDD 2002
Automated Learning of Decision Rules for Text Categorization · ACM Trans. Inf. Syst. 1994
Data mining › predictive modeling
classification
0.132001
Solving regression problems with rule-based ensemble classifiers · KDD 2001
Lightweight Rule Induction · ICML 2000
Small Sample Decision tree Pruning · ICML 1994
Data mining › predictive modeling › classification
rule induction
0.122003
Knowledge-based data mining · KDD 2003
Lightweight Rule Induction · ICML 2000
Data mining › predictive modeling › classification
ensemble learning
0.012001
Solving regression problems with rule-based ensemble classifiers · KDD 2001
Data mining › predictive modeling
regression
0.012001
Solving regression problems with rule-based ensemble classifiers · KDD 2001
Knowledge, reasoning and agents › Knowledge representation and reasoning
expert systems
0.072003
Knowledge-based data mining · KDD 2003
An Approach to Expert Control of Interactive Software Systems · IEEE Trans. Pattern Anal. Mach. Intell. 1985
Using Empirical Analysis to Refine Expert System Knowledge Bases · Artif. Intell. 1984
Data mining › dimensionality reduction
feature extraction
0.011995
Feature Extraction for Massive Data Mining · KDD 1995
Machine learning › Kernel, tree and ensemble methods
decision tree
0.011994
Decision Tree Pruning: Biased or Optimal? · AAAI 1994
Machine learning › Kernel, tree and ensemble methods › decision tree learning
decision tree pruning
0.011994
Decision Tree Pruning: Biased or Optimal? · AAAI 1994
Data mining › predictive modeling › classification
decision tree learning
0.011994
Small Sample Decision tree Pruning · ICML 1994
Data mining › predictive modeling › classification › decision tree learning
decision tree pruning
0.011994
Small Sample Decision tree Pruning · ICML 1994
Information retrieval
retrieval models
0.011994
Automated Learning of Decision Rules for Text Categorization · ACM Trans. Inf. Syst. 1994
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge base
knowledge base refinement
0.031988
Automatic Knowledge Base Refinement for Classification Systems · Artif. Intell. 1988
SEEK2: A Generalized Approach to Automatic Knowledge Base Refinement · IJCAI 1985
Using Empirical Analysis to Refine Expert System Knowledge Bases · Artif. Intell. 1984
Knowledge, reasoning and agents › Knowledge representation and reasoning
rule learning
0.021994
Reduced Complexity Rule Induction · IJCAI 1991
Automated Learning of Decision Rules for Text Categorization · ACM Trans. Inf. Syst. 1994
Information retrieval › evaluation
benchmark
0.012002
Experiments in high-dimensional text categorization · SIGIR 2002
Information retrieval
evaluation
0.012002
Experiments in high-dimensional text categorization · SIGIR 2002
Bioinformatics and computational biology
protein structure prediction
0.011993
Transmembrane Segment Prediction from Protein Sequence Data · ISMB 1993
Bioinformatics and computational biology › protein structure prediction › membrane protein structure prediction › transmembrane topology prediction
transmembrane helix prediction
0.011993
Transmembrane Segment Prediction from Protein Sequence Data · ISMB 1993
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge-based systems › rule-based systems
production rules
0.021990
Maximizing the Predictive Value of Production Rules · Artif. Intell. 1990
Learning Production Rules for Consultation Systems · IJCAI 1979
Machine learning › Learning theory › classification
classifier evaluation
0.011991
Small Sample Error Rate Estimation for k-NN Classifiers · IEEE Trans. Pattern Anal. Mach. Intell. 1991
Machine learning › Learning theory › model selection
cross-validation
0.011991
Small Sample Error Rate Estimation for k-NN Classifiers · IEEE Trans. Pattern Anal. Mach. Intell. 1991
Machine learning › Learning theory › statistical estimation › risk estimation
error rate estimation
0.011991
Small Sample Error Rate Estimation for k-NN Classifiers · IEEE Trans. Pattern Anal. Mach. Intell. 1991
Machine learning › Learning theory
statistical learning theory
0.011991
Small Sample Error Rate Estimation for k-NN Classifiers · IEEE Trans. Pattern Anal. Mach. Intell. 1991
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge-based systems
rule-based systems
0.011990
Maximizing the Predictive Value of Production Rules · Artif. Intell. 1990
Knowledge, reasoning and agents › Knowledge representation and reasoning › expert systems
uncertainty in expert systems
0.011990
Maximizing the Predictive Value of Production Rules · Artif. Intell. 1990
Knowledge, reasoning and agents › Knowledge representation and reasoning › decision theory
decision rule
0.011987
Optimizing the Predictive Value of Diagnostic Decision Rules · AAAI 1987
Data mining › big data analytics
large-scale data mining
0.011995
Feature Extraction for Massive Data Mining · KDD 1995
Machine learning › Learning theory › statistical learning theory
bias-variance tradeoff
0.011994
Decision Tree Pruning: Biased or Optimal? · AAAI 1994
Natural language and speech › Information extraction and text analysis
text classification
0.011994
Towards Language Independent Automated Learning of Text Categorisation Models · SIGIR 1994
Medical and health informatics
clinical decision support
0.031981
A Model-Based Consultation System for the Long-Term Management of Glaucoma · IJCAI 1977
A Precedence Scheme for Selection and Explanation of Therapies · IJCAI 1981
A Model-Based Method for Computer-Aided Medical Decision-Making · Artif. Intell. 1978

Methods — techniques the papers use, named apart from their topics

statistical learning · 0.2missing data imputation · 0.2decision tree · 0.1decision rules · 0.1rule induction · 0.1text analysis · 0.1sparse feature representation · 0.0linear classifier · 0.0k-means clustering · 0.0decision-rule ensembles · 0.0automated learning · 0.0rule-based induction · 0.0pruning · 0.0feature selection · 0.0decision tree pruning · 0.0monte carlo simulation · 0.0leaving-one-out · 0.0bootstrap · 0.0
YearPublicationVenuePosition
2015 Managing healthcare costs by peer-group modeling
Sholom M. Weiss, Casimir A. Kulikowski, Robert S. Galen, Peder A. Olsen, Ramesh Natarajan
Appl. Intell.1
2014 Graphical Models for Identifying Fraud and Waste in Healthcare Claims
abstract
We describe graphical model based methods for analyzing prescription and medical claims data in order to identify fraud and waste. Our approach draws on ideas from speech recognition and language modeling to identify patients, doctors and pharmacies whose prescription encounters show significant departure from normative behavior. We have analyzed claims data from a large healthcare provider, consisting of over 53 million individual prescription claims in the calendar year 2011.
Peder A. Olsen, Ramesh Natarajan, Sholom M. Weiss
SDM3
2013 Improving quality control by early prediction of manufacturing outcomes
abstract
We describe methods for continual prediction of manufactured product quality prior to final testing. In our most expansive modeling approach, an estimated final characteristic of a product is updated after each manufacturing operation. Our initial application is for the manufacture of microprocessors, and we predict final microprocessor speed. Using these predictions, early corrective manufacturing actions may be taken to increase the speed of expected slow wafers (a collection of microprocessors) or reduce the speed of fast wafers. Such predictions may also be used to initiate corrective supply chain management actions. Developing statistical learning models for this task has many complicating factors: (a) a temporally unstable population (b) missing data that is a result of sparsely sampled measurements and (c) relatively few available measurements prior to corrective action opportunities. In a real manufacturing pilot application, our automated models selected 125 fast wafers in real-time. As predicted, those wafers were significantly faster than average. During manufacture, downstream corrective processing restored 25 nominally unacceptable wafers to normal operation.
Sholom M. Weiss, Amit Dhurandhar, Robert J. Baseman
KDD1
2010 Rule-based data mining for yield improvement in semiconductor manufacturing
Sholom M. Weiss, Robert J. Baseman, Fateh Tipu, Christopher N. Collins, William A. Davies, Raminderpal Singh, John W. Hopkins
Appl. Intell.1
2008 Estimating Sales Opportunity Using Similarity-Based Methods
Sholom M. Weiss, Nitin Indurkhya
ECML/PKDD (2)1
2004 Text categorization for a comprehensive time-dependent benchmark
Fred J. Damerau, Tong Zhang 0001, Sholom M. Weiss, Nitin Indurkhya
Inf. Process. Manag.3
2003 Knowledge-based data mining
abstract
We describe techniques for combining two types of knowledge systems: expert and machine learning. Both the expert system and the learning system represent information by logical decision rules or trees. Unlike the classical views of knowledge-base evaluation or refinement, our view accepts the contents of the knowledge base as completely correct. The knowledge base and the results of its stored cases will provide direction for the discovery of new relationships in the form of newly induced decision rules. An expert system called SEAS was built to discover sales leads for computer products and solutions. The system interviews executives by asking questions, and based on the responses, recommends products that may improve a business' operations. Leveraging this expert system, we record the results of the interviews and the program's recommendations. The very same data stored by the expert system is used to find new predictive rules. Among the potential advantages of this approach are (a) the capability to spot new sales trends and (b) the substitution of less expensive probabilistic rules that use database data instead of interviews.
Sholom M. Weiss, Stephen J. Buckley, Shubir Kapoor, Søren Damgaard
KDD1
2002 A system for real-time competitive market intelligence
abstract
A method is described for real-time market intelligence and competitive analysis. News stories are collected online for a designated group of companies. The goal is to detect critical differences in the text written about a company versus the text for its competitors. A solution is found by mapping the task into a non-stationary text categorization model. The overall design consists of the following components: (a) a real-time crawler that monitors newswires for stories about the competitors (b) a conditional document retriever that selects only those documents that meet the indicated conditions (c) text analysis techniques that convert the documents to a numerical format (d) rule induction methods for finding patterns in data (e) presentation techniques for displaying results. The method is extended to combine text with numerical measures, such as those based on stock prices and market capitalizations, that allow for more objective evaluations and projections.
Sholom M. Weiss, Naval K. Verma
KDD1
2002 Experiments in high-dimensional text categorization
abstract
We present results for automated text categorization of the Reuters-810000 collection of news stories. Our experiments use the entire one-year collection of 810,000 stories and the entire subject index. We divide the data into monthly groups and provide an initial benchmark of text categorization performance on the complete collection. Experimental results show that efficient sparse-feature implementations of linear methods and decision trees, using a global unstemmed dictionary, can readily handle applications of this size. Predictive performance is approximately as strong as the best results for the much smaller older Reuters collections. Detailed results are provided over time periods. It is shown that a smaller time horizon does not diminish predictive quality, implying reduced demands for retraining when sample size is large.
Fred J. Damerau, Tong Zhang 0001, Sholom M. Weiss, Nitin Indurkhya
SIGIR3
2001 Solving regression problems with rule-based ensemble classifiers
abstract
We describe a lightweight learning method that induces an ensemble of decision-rule solutions for regression problems. Instead of direct prediction of a continuous output variable, the method discretizes the variable by k-means clustering and solves the resultant classification problem. Predictions on new examples are made by averaging the mean values of classes with votes that are close in number to the most likely class. We provide experimental evidence that this indirect approach can often yield strong results for many applications, generally outperforming direct approaches such as regression trees and rivaling bagged regression trees.
Nitin Indurkhya, Sholom M. Weiss
KDD2
2001 Lightweight Collaborative Filtering Method for Binary-Encoded Data
Sholom M. Weiss, Nitin Indurkhya
PKDD1
2001 Advances in predictive models for data mining
Se June Hong, Sholom M. Weiss
Pattern Recognit. Lett.2
2000 Lightweight Rule Induction
Sholom M. Weiss, Nitin Indurkhya
ICML1
2000 Leightweight Document Clustering
Sholom M. Weiss, Brian F. White, Chidanand Apté
PKDD1
1998 Estimating Performance Gains for Voted Decision Trees
abstract
Decision tree induction is a prominent learning method, typically yielding quick results with competitive predictive performance. However, it is not unusual to find other automated learning methods that exceed the predictive performance of a decision tree on the same application. To achieve near-optimal classification results, resampling techniques can be employed to generate multiple decision-tree solutions. These decision trees are individually applied and their answers voted. The potential for exceptionally strong performance is counterbalanced by the substantial increase in computing time to induce many decision trees. We describe estimators of predictive performance for voted decision trees induced from bootstrap (bagged) or adaptive (boosted) resampling. The estimates are found by examining the performance of a single tree and its pruned subtrees over a single, training set and a large test set. Using publicly available collections of data, we show that these estimates are usually quite accurate, with occasional weaker estimates. The great advantage of these estimates is that they reveal the predictive potential of voted decision trees prior to applying expensive computational procedures.
Nitin Indurkhya, Sholom M. Weiss
Intell. Data Anal.2
1997 Data mining with decision trees and decision rules
Chidanand Apté, Sholom M. Weiss
Future Gener. Comput. Syst.2
1996 Selecting the Right-Size Model for Prediction
Sholom M. Weiss, Nitin Indurkhya
Appl. Intell.1
1995 Using Case Data to Improve on Rule-based Function Approximation
Nitin Indurkhya, Sholom M. Weiss
ICCBR2
1995 Feature Extraction for Massive Data Mining
Raguram Sasisekharan, Sholom M. Weiss
KDD3
1995 Rule-based Machine Learning Methods for Functional Prediction
abstract
We describe a machine learning method for predicting the value of a real-valued function, given the values of multiple input variables. The method induces solutions from samples in the form of ordered disjunctive normal form (DNF) decision rules. A central objective of the method and representation is the induction of compact, easily interpretable solutions. This rule-based decision model can be extended to search efficiently for similar cases prior to approximating function values. Experimental results on real-world data demonstrate that the new techniques are competitive with existing machine learning and statistical methods and can sometimes yield superior regression performance.
Sholom M. Weiss, Nitin Indurkhya
J. Artif. Intell. Res.1
1994 Decision Tree Pruning: Biased or Optimal?
Sholom M. Weiss, Nitin Indurkhya
AAAI1
1994 Small Sample Decision tree Pruning
Sholom M. Weiss, Nitin Indurkhya
ICML1
1994 Towards Language Independent Automated Learning of Text Categorisation Models
Chidanand Apté, Fred J. Damerau, Sholom M. Weiss
SIGIR3
1994 Case studies in high-dimensional classification
Chidanand Apté, Raguram Sasisekharan, Sholom M. Weiss
Appl. Intell.4
1994 Guest editors' introduction
Sholom M. Weiss, Nitin Indurkhya
Appl. Intell.1
1994 Automated Learning of Decision Rules for Text Categorization
abstract
We describe the results of extensive experiments using optimized rule-based induction methods on large document collections. The goal of these methods is to discover automatically classification patterns that can be used for general document categorization or personalized filtering of free text. Previous reports indicate that human-engineered rule-based systems, requiring many man-years of developmental efforts, have been successfully built to “read” documents and assign topics to them. We show that machine-generated decision rules appear comparable to human performance, while using the identical rule-based representation. In comparison with other machine-learning techniques, results on a key benchmark from the Reuters collection show a large gain in performance, from a previously reported 67% recall/precision breakeven point to 80.5%. In the context of a very high-dimensional feature space, several methodological alternatives are examined, including universal versus local dictionaries, and binary versus frequency-related features.
Chidanand Apté, Fred J. Damerau, Sholom M. Weiss
ACM Trans. Inf. Syst.3
1993 Rule-Based Regression
Sholom M. Weiss, Nitin Indurkhya
IJCAI1
1993 Transmembrane Segment Prediction from Protein Sequence Data
Sholom M. Weiss, Dawn M. Cohen, Nitin Indurkhya
ISMB1
1992 Heuristic configuration of single hidden-layer feed-forward neural networks
Nitin Indurkhya, Sholom M. Weiss
Appl. Intell.2
1991 Reduced Complexity Rule Induction
Sholom M. Weiss, Nitin Indurkhya
IJCAI1
1991 Iterative rule induction methods
Nitin Indurkhya, Sholom M. Weiss
Appl. Intell.2
1991 Small Sample Error Rate Estimation for k-NN Classifiers
abstract
Small sample error rate estimators for nearest-neighbor classifiers are examined and contrasted with the same estimators for three-nearest-neighbor classifiers. The performance of the bootstrap estimators, e0 and 0.632B, is considered relative to leaving-one-out and other cross-validation estimators. Monte Carlo simulations are used to measure the performance of the error-rate estimators. The experimental results are compared to previously reported simulations for nearest-neighbor classifiers and alternative classifiers. It is shown that each of the estimators has strengths and weaknesses for varying apparent and true error-rate situations. A combined estimator that corrects the leaving-one-out estimator (by combining bootstrap and cross-validation estimators) gives strong results over a broad range of situations.>
Sholom M. Weiss
IEEE Trans. Pattern Anal. Mach. Intell.1
1990 Maximizing the Predictive Value of Production Rules
Sholom M. Weiss, Robert S. Galen, Prasad Tadepalli
Artif. Intell.1
1989 An Empirical Comparison of Pattern Recognition, Neural Nets, and Machine Learning Classification Methods
Sholom M. Weiss, Ioannis Kapouleas
IJCAI1
1989 Models for measuring performance of medical expert systems
Nitin Indurkhya, Sholom M. Weiss
Artif. Intell. Medicine2
1988 Automatic Knowledge Base Refinement for Classification Systems
Allen Ginsberg, Sholom M. Weiss, Peter Politakis
Artif. Intell.2
1987 Optimizing the Predictive Value of Diagnostic Decision Rules
Sholom M. Weiss, Robert S. Galen, Prasad Tadepalli
AAAI1
1985 SEEK2: A Generalized Approach to Automatic Knowledge Base Refinement
Allen Ginsberg, Sholom M. Weiss, Peter Politakis
IJCAI2
1985 An Approach to Expert Control of Interactive Software Systems
abstract
Expert problem-solving strategies in many domains require the use of detailed mathematical techniques coupled with experiential knowledge about how and when to use the appropriate techniques. In many of these domains, such techniques are made available to experts in large software packages. In attempting to build expert systems for these domains, we wish to make use of these packages, and are therefore faced with an important problem: how to integrate the existing software, and knowledge about its use, into a practical expert system. The expert knowledge is used, in dynamic selection and interpretation of appropriate programs and parameters, to reach a successful goal in the problem solving. We describe the framework of a hybrid expert system for representing problem-solving knowledge in these domains. This hybrid system may be characterized as consisting of a production system and mathematical methods. The software package is reorganized as necessary to map it into the mathematical-method representation of a hybrid system. This approach has evolved out of an effort to build an expert system for performing well-log analysis, ELAS (expert log analysis system).
Chidanand Apté, Sholom M. Weiss
IEEE Trans. Pattern Anal. Mach. Intell.2
1984 Using Empirical Analysis to Refine Expert System Knowledge Bases
Peter Politakis, Sholom M. Weiss
Artif. Intell.2
1982 Building Expert Systems for Controlling Complex Programs
Sholom M. Weiss, Casimir A. Kulikowski, Chidanand Apté, Michael Uschold, Jay Patchett, Robert Brigham, Belynda Spitzer
AAAI1
1981 A Precedence Scheme for Selection and Explanation of Therapies
John K. Kastmer, Sholom M. Weiss
IJCAI2
1981 Developing Microprocessor Based Expert Models for Instrument Interpretation
Sholom M. Weiss, Casimir A. Kulikowski, Robert S. Galen
IJCAI1
1979 EXPERT: A System for Developing Consultation Models
Sholom M. Weiss, Casimir A. Kulikowski
IJCAI1
1979 Learning Production Rules for Consultation Systems
Sholom M. Weiss, Casimir A. Kulikowski, Bernard Nudel
IJCAI1
1978 A Model-Based Method for Computer-Aided Medical Decision-Making
Sholom M. Weiss, Casimir A. Kulikowski, Saul Amarel, Aran Safir
Artif. Intell.1
1977 A Model-Based Consultation System for the Long-Term Management of Glaucoma
Sholom M. Weiss, Casimir A. Kulikowski, Aran Safir
IJCAI1