Jerome H. Friedman

dblp:92/5390 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
1since 2021 · last 2022
0000-0001-5968-8901ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 1 since 2021Systems, architecture and hardware · 4 · 3 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorTheory of computation · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Kernel, tree and ensemble methods · 69% Transfer learning and domain adaptation · 20% Optimization for machine learning · 10%
Theoretical computer science
5 papers
Mathematical optimization · 95% Graph algorithms and graph theory · 3% Algorithms and data structures · 2%

Topics — the 11 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Kernel, tree and ensemble methods
gradient boosting
0.612022
Representational Gradient Boosting: Backpropagation in the Space of Functions · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Machine learning › Transfer learning and domain adaptation
meta-learning
0.212022
Representational Gradient Boosting: Backpropagation in the Space of Functions · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Machine learning › Optimization for machine learning
coordinate descent
0.112008
Regularization paths and coordinate descent · KDD 2008
Mathematical optimization
regularization
0.112008
Regularization paths and coordinate descent · KDD 2008
Graph algorithms and graph theory
graph algorithms
0.011978
Fast Algorithms for Constructing Minimal Spanning Trees in Coordinate Spaces · IEEE Trans. Computers 1978
Graph algorithms and graph theory › spanning tree
minimum spanning tree
0.011978
Fast Algorithms for Constructing Minimal Spanning Trees in Coordinate Spaces · IEEE Trans. Computers 1978
Algorithms and data structures › similarity search
nearest neighbor search
0.021977
An Algorithm for Finding Nearest Neighbors · IEEE Trans. Computers 1975
A Recursive Partitioning Decision Rule for Nonparametric Classification · IEEE Trans. Computers 1977
Machine learning › Learning theory › classification
nonparametric classification
0.011977
A Recursive Partitioning Decision Rule for Nonparametric Classification · IEEE Trans. Computers 1977
Machine learning › Kernel, tree and ensemble methods
recursive partitioning
0.011977
A Recursive Partitioning Decision Rule for Nonparametric Classification · IEEE Trans. Computers 1977
Algorithms and data structures › similarity search › nearest neighbor search
k-nearest neighbors
0.011975
An Algorithm for Finding Nearest Neighbors · IEEE Trans. Computers 1975
Machine learning › Optimization for machine learning
projection pursuit
0.011974
A Projection Pursuit Algorithm for Exploratory Data Analysis · IEEE Trans. Computers 1974

Methods — techniques the papers use, named apart from their topics

stacking · 0.6gradient boosting · 0.6backpropagation · 0.6lasso · 0.2coordinate descent · 0.2recursive partitioning · 0.0projection pursuit · 0.0bayes risk analysis · 0.0distance calculation pruning · 0.0computational geometry · 0.0
YearPublicationVenuePosition
2022 Representational Gradient Boosting: Backpropagation in the Space of Functions
abstract
The estimation of nested functions (i.e., functions of functions) is one of the central reasons for the success and popularity of machine learning. Today, artificial neural networks are the predominant class of algorithms in this area, known as representational learning. Here, we introduce Representational Gradient Boosting (RGB), a nonparametric algorithm that estimates functions with multi-layer architectures obtained using backpropagation in the space of functions. RGB does not need to assume a functional form in the nodes or output (e.g., linear models or rectified linear units), but rather estimates these transformations. RGB can be seen as an optimized stacking procedure where a meta algorithm learns how to combine different classes of functions (e.g., Neural Networks (NN) and Gradient Boosting (GB)), while building and optimizing them jointly in an attempt to compensate each other's weaknesses. This highlights a stark difference with current approaches to meta-learning that combine models only after they have been built independently. We showed that providing optimized stacking is one of the main advantages of RGB over current approaches. Additionally, due to the nested nature of RGB we also showed how it improves over GB in problems that have several high-order interactions. Finally, we investigate both theoretically and in practice the problem of recovering nested functions and the value of prior knowledge.
Gilmer Valdes, Jerome H. Friedman, Efstathios D. Gennatas
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Intruders pattern identification
abstract
This paper considers the problem of intrusion detection in information systems as a classification problem. In particular the case of masquerader is treated. This kind of intrusion is one of the more difficult to discover because it may attack already open user sessions. Moreover, this problem is complex because of the large variability of user models and the lack of available data for the learning purpose. Here, flexible and robust similarity measures, suitable also for non-numeric data, are defined, they will be incorporated on a one-class training K N N and compared with several classification methods proposed in the literature using the Masquerading User Data set (www.schonlau.net) representing users and intruders on an UNIX system.
Vito Di Gesù, Giosuè Lo Bosco, Jerome H. Friedman
ICPR3
2008 Regularization paths and coordinate descent
abstract
In a statistical world faced with an explosion of data, regularization has become an important ingredient. In a wide variety of problems we have many more input features than observations, and the lasso penalty and its hybrids have become increasingly useful for both feature selection and regularization. This talk presents some effective algorithms based on coordinate descent for fitting large scale regularization paths for a variety of problems.
Trevor J. Hastie, Jerome H. Friedman, Robert Tibshirani
KDD2
2003 Note on "Comparison of Model Selection for Regression" by Vladimir Cherkassky and Yunqian Ma
abstract
While Cherkassky and Ma (2003) raise some interesting issues in comparing techniques for model selection, their article appears to be written largely in protest of comparisons made in our book, Elements of Statistical Learning (2001). Cherkassky and Ma feel that we falsely represented the structural risk minimization (SRM) method, which they defend strongly here. In a two-page section of our book (pp. 212-213), we made an honest attempt to compare the SRM method with two related techniques, Aikaike information criterion (AIC) and Bayesian information criterion (BIC). Apparently, we did not apply SRM in the optimal way. We are also accused of using contrived examples, designed to make SRM look bad. Alas, we did introduce some careless errors in our original simulation--errors that were corrected in the second and subsequent printings. Some of these errors were pointed out to us by Cherkassky and Ma (we supplied them with our source code), and as a result we replaced the assessment "SRM performs poorly overall" with a more moderate "the performance of SRM is mixed" (p. 212).
Trevor J. Hastie, Robert Tibshirani, Jerome H. Friedman
Neural Comput.3
1997 On Bias, Variance, 0/1-Loss, and the Curse-of-Dimensionality
Jerome H. Friedman
Data Min. Knowl. Discov.1
1990 Adaptive Spline Networks
Jerome H. Friedman
NIPS1
1981 A Nested Partitioning Procedure for Numerical Multiple Integration
abstract
An algorlthrs is presented for adaptively partitioning a multidirsenslonal coorchnate space on the basis of optimization of a scalar functmn of the coordinates.The goal is to construct a set of hyperrectangular regions, such that the variation of function values within each region is small.These regions are then used as the basis for a stratified samplmg estlrsate of the defirste integral of the function.
Jerome H. Friedman, Margaret H. Wright
ACM Trans. Math. Softw.1
1978 Fast Algorithms for Constructing Minimal Spanning Trees in Coordinate Spaces
abstract
Algorithms are presented that construct the shortest connecting network, or minimal spanning tree (MST), of N points embedded in k-dimensional coordinate space. These algorithms take advantage of the geometry of such spaces to substantially reduce the computation from that required to construct MST's of more general graphs. An algorithm is also presented that constructs a spanning tree that is very nearly minimal with computation proportional to N log N for all k.
Jon Louis Bentley, Jerome H. Friedman
IEEE Trans. Computers2
1977 A Recursive Partitioning Decision Rule for Nonparametric Classification
abstract
A new criterion for deriving a recursive partitioning decision rule for nonparametric classification is presented. The criterion is both conceptually and computationally simple, and can be shown to have strong statistical merit. The resulting decision rule is asymptotically Bayes' risk efficient. The notion of adaptively generated features is introduced and methods are presented for dealing with missing features in both training and test vectors.
Jerome H. Friedman
IEEE Trans. Computers1
1977 An Algorithm for Finding Best Matches in Logarithmic Expected Time
abstract
An algorithm and data structure are presented for searching a file containing N records, each described by k real valued keys, for the m closest matches or nearest neighbors to a given query record. The computation required to organize the file is proportional to kNlogN. The expected number of records examined in each search is independent of the file size. The expected computation to perform each search is proportional-to 1ogN. Empirical evidence suggests that except for very small files, this algorithm is considerably faster than other methods.
Jerome H. Friedman, Jon Louis Bentley, Raphael A. Finkel
ACM Trans. Math. Softw.1
1975 An Algorithm for Finding Nearest Neighbors
abstract
An algorithm that finds the k nearest neighbors of a point, from a sample of size N in a d-dimensional space, with an expected number of distance calculations is described, its properties examined, and the validity of the estimate verified with simulated data.
Jerome H. Friedman, Forest Baskett, Leonard J. Shustek
IEEE Trans. Computers1
1974 A Projection Pursuit Algorithm for Exploratory Data Analysis
abstract
An algorithm for the analysis of multivariate data is presented and is discussed in terms of specific examples. The algorithm seeks to find one-and two-dimensional linear projections of multivariate data that are relatively highly revealing.
Jerome H. Friedman, John W. Tukey
IEEE Trans. Computers1