Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Rie Johnson

dblp:66/1605 · DBLP profile ↗
← Back
13ranked-venue papers
12as first author
2since 2021 · last 2023
0000-0001-7979-3856ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 11 first-author · 2 since 2021Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Learning theory · 31% Information extraction and text analysis · 17% Generative modeling · 14%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 16 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
generative adversarial network
0.822021
A Framework of Composite Functional Gradient Methods for Generative Adversarial Models · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Composite Functional Gradient Learning of Generative Adversarial Models · ICML 2018
Natural language and speech › Information extraction and text analysis
text classification
0.832017
Deep Pyramid Convolutional Neural Networks for Text Categorization · ACL (1) 2017
Supervised and Semi-Supervised Text Categorization using LSTM for Region Embeddings · ICML 2016
Semi-supervised Convolutional Neural Networks for Text Categorization via Region Embedding · NIPS 2015
Machine learning › Learning theory › generalization error
generalization gap
0.712023
Inconsistency, Instability, and Generalization Gap of Deep Neural Network Training · NeurIPS 2023
Machine learning › Learning theory › generalization
stability and generalization
0.712023
Inconsistency, Instability, and Generalization Gap of Deep Neural Network Training · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation learning › embedding learning › image embedding
region embedding
0.522016
Supervised and Semi-Supervised Text Categorization using LSTM for Region Embeddings · ICML 2016
Semi-supervised Convolutional Neural Networks for Text Categorization via Region Embedding · NIPS 2015
Machine learning › Optimization for machine learning
non-convex optimization
0.412020
Guided Learning of Nonconvex Models through Successive Functional Gradient Optimization · ICML 2020
Machine learning › Deep learning architectures and training
convolutional neural network
0.312017
Deep Pyramid Convolutional Neural Networks for Text Categorization · ACL (1) 2017
Natural language and speech › Information extraction and text analysis › sentiment analysis
sentiment classification
0.312017
Deep Pyramid Convolutional Neural Networks for Text Categorization · ACL (1) 2017
Machine learning › Optimization for machine learning
stochastic gradient descent
0.212013
Accelerating Stochastic Gradient Descent using Predictive Variance Reduction · NIPS 2013
Mathematical optimization › stochastic optimization
variance reduction
0.212013
Accelerating Stochastic Gradient Descent using Predictive Variance Reduction · NIPS 2013
Machine learning › Learning paradigms
semi-supervised learning
0.222016
Graph-Based Semi-Supervised Learning and Spectral Kernel Design · IEEE Trans. Inf. Theory 2008
Supervised and Semi-Supervised Text Categorization using LSTM for Region Embeddings · ICML 2016
Machine learning › Learning paradigms › semi-supervised learning
graph-based semi-supervised learning
0.222008
Graph-Based Semi-Supervised Learning and Spectral Kernel Design · IEEE Trans. Inf. Theory 2008
On the Effectiveness of Laplacian Normalization for Graph Semi-supervised Learning · J. Mach. Learn. Res. 2007
Machine learning › Probabilistic and Bayesian machine learning
divergence minimization
0.112018
Composite Functional Gradient Learning of Generative Adversarial Models · ICML 2018
Machine learning › Kernel, tree and ensemble methods
spectral kernel design
0.112008
Graph-Based Semi-Supervised Learning and Spectral Kernel Design · IEEE Trans. Inf. Theory 2008
Machine learning › Graph learning
graph representation learning
0.112007
On the Effectiveness of Laplacian Normalization for Graph Semi-supervised Learning · J. Mach. Learn. Res. 2007
Machine learning › Learning theory
generalization bounds
0.012008
Graph-Based Semi-Supervised Learning and Spectral Kernel Design · IEEE Trans. Inf. Theory 2008

Methods — techniques the papers use, named apart from their topics

functional gradient · 0.8convolutional neural network · 0.8ensemble · 0.7co-distillation · 0.7minimax optimization · 0.5composite functional gradient learning · 0.5mirror descent · 0.4KL divergence · 0.3region embedding · 0.2LSTM · 0.2variance reduction · 0.2convergence analysis · 0.2
YearPublicationVenuePosition
2023 Inconsistency, Instability, and Generalization Gap of Deep Neural Network Training
abstract
As deep neural networks are highly expressive, it is important to find solutions with small generalization gap (the difference between the performance on the training data and unseen data). Focusing on the stochastic nature of training, we first present a theoretical analysis in which the bound of generalization gap depends on what we call inconsistency and instability of model outputs, which can be estimated on unlabeled data. Our empirical study based on this analysis shows that instability and inconsistency are strongly predictive of generalization gap in various settings. In particular, our finding indicates that inconsistency is a more reliable indicator of generalization gap than the sharpness of the loss landscape. Furthermore, we show that algorithmic reduction of inconsistency leads to superior performance. The results also provide a theoretical basis for existing methods such as co-distillation and ensemble.
Rie Johnson, Tong Zhang 0001
NeurIPS1
2021 A Framework of Composite Functional Gradient Methods for Generative Adversarial Models
abstract
Generative adversarial networks (GAN) are trained through a minimax game between a generator and a discriminator to generate data that mimics observations. While being widely used, GAN training is known to be empirically unstable. This paper presents a new theory for generative adversarial methods that does not rely on the traditional minimax formulation. Our theory shows that with a strong discriminator, a good generator can be obtained by composite functional gradient learning, so that several distance measures (including the KL divergence and the JS divergence) between the probability distributions of real data and generated data are simultaneously improved after each functional gradient step until converging to zero. This new point of view leads to stable procedures for training generative models. It also gives a new theoretical insight into the original GAN. Empirical results on image generation show the effectiveness of our new method.
Rie Johnson, Tong Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 Guided Learning of Nonconvex Models through Successive Functional Gradient Optimization
abstract
This paper presents a framework of successive functional gradient optimization for training nonconvex models such as neural networks, where training is driven by mirror descent in a function space. We provide a theoretical analysis and empirical study of the training method derived from this framework. It is shown that the method leads to better performance than that of standard training techniques.
Rie Johnson, Tong Zhang 0001
ICML1
2018 Composite Functional Gradient Learning of Generative Adversarial Models
abstract
This paper first presents a theory for generative adversarial methods that does not rely on the traditional minimax formulation. It shows that with a strong discriminator, a good generator can be learned so that the KL divergence between the distributions of real data and generated data improves after each functional gradient step until it converges to zero. Based on the theory, we propose a new stable generative adversarial method. A theoretical insight into the original GAN from this new viewpoint is also provided. The experiments on image generation show the effectiveness of our new method.
Rie Johnson, Tong Zhang 0001
ICML1
2017 Deep Pyramid Convolutional Neural Networks for Text Categorization
abstract
This paper proposes a low-complexity word-level deep convolutional neural network (CNN) architecture for text categorization that can efficiently represent longrange associations in text.In the literature, several deep and complex neural networks have been proposed for this task, assuming availability of relatively large amounts of training data.However, the associated computational complexity increases as the networks go deeper, which poses serious challenges in practical applications.Moreover, it was shown recently that shallow word-level CNNs are more accurate and much faster than the state-of-the-art very deep nets such as character-level CNNs even in the setting of large training data.Motivated by these findings, we carefully studied deepening of word-level CNNs to capture global representations of text, and found a simple network architecture with which the best accuracy can be obtained by increasing the network depth without increasing computational cost by much.We call it deep pyramid CNN.The proposed model with 15 weight layers outperforms the previous best models on six benchmark datasets for sentiment classification and topic categorization.
Rie Johnson, Tong Zhang 0001
ACL (1)1
2016 Supervised and Semi-Supervised Text Categorization using LSTM for Region Embeddings
abstract
One-hot CNN (convolutional neural network) has been shown to be effective for text categorization (Johnson & Zhang, 2015). We view it as a special case of a general framework which jointly trains a linear model with a non-linear feature generator consisting of ‘text region embedding + pooling’. Under this framework, we explore a more sophisticated region embedding method using Long Short-Term Memory (LSTM). LSTM can embed text regions of variable (and possibly large) sizes, whereas the region size needs to be fixed in a CNN. We seek effective and efficient use of LSTM for this purpose in the supervised and semi-supervised settings. The best results were obtained by combining region embeddings in the form of LSTM and convolution layers trained on unlabeled data. The results indicate that on this task, embeddings of text regions, which can convey complex concepts, are more useful than embeddings of single words in isolation. We report performances exceeding the previous best results on four benchmark datasets.
Rie Johnson, Tong Zhang 0001
ICML1
2015 Effective Use of Word Order for Text Categorization with Convolutional Neural Networks
abstract
Convolutional neural network (CNN) is a neural network that can make use of the internal structure of data such as the 2D structure of image data. This paper studies CNN on text categorization to exploit the 1D structure (namely, word order) of text data for accurate prediction. Instead of using low-dimensional word vectors as input as is often done, we directly apply CNN to high-dimensional text data, which leads to directly learning embedding of small text regions for use in classification. In addition to a straightforward adaptation of CNN from image to text, a simple but new variation which employs bag-of-word conversion in the convolution layer is proposed. An extension to combine multiple convolution layers is also explored for higher accuracy. The experiments demonstrate the effectiveness of our approach in comparison with state-of-the-art methods.
Rie Johnson, Tong Zhang 0001
HLT-NAACL1
2015 Semi-supervised Convolutional Neural Networks for Text Categorization via Region Embedding
abstract
This paper presents a new semi-supervised framework with convolutional neural networks (CNNs) for text categorization. Unlike the previous approaches that rely on word embeddings, our method learns embeddings of small text regions from unlabeled data for integration into a supervised CNN. The proposed scheme for embedding learning is based on the idea of two-view semi-supervised learning, which is intended to be useful for the task of interest even though the training is done on unlabeled data. Our models achieve better results than previous approaches on sentiment classification and topic classification tasks.
Rie Johnson, Tong Zhang 0001
NIPS1
2014 Learning Nonlinear Functions Using Regularized Greedy Forest
abstract
We consider the problem of learning a forest of nonlinear decision rules with general loss functions. The standard methods employ boosted decision trees such as Adaboost for exponential loss and Friedman's gradient boosting for general loss. In contrast to these traditional boosting algorithms that treat a tree learner as a black box, the method we propose directly learns decision forests via fully-corrective regularized greedy search using the underlying forest structure. Our method achieves higher accuracy and smaller models than gradient boosting on many of the datasets we have tested on.
Rie Johnson, Tong Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2013 Accelerating Stochastic Gradient Descent using Predictive Variance Reduction
abstract
Stochastic gradient descent is popular for large scale optimization but has slow convergence asymptotically due to the inherent variance. To remedy this problem, we introduce an explicit variance reduction method for stochastic gradient descent which we call stochastic variance reduced gradient (SVRG). For smooth and strongly convex functions, we prove that this method enjoys the same fast convergence rate as those of stochastic dual coordinate ascent (SDCA) and Stochastic Average Gradient (SAG). However, our analysis is significantly simpler and more intuitive. Moreover, unlike SDCA or SAG, our method does not require the storage of gradients, and thus is more easily applicable to complex problems such as some structured prediction problems and neural network learning.
Rie Johnson, Tong Zhang 0001
NIPS1
2008 Word sense disambiguation across two domains: Biomedical literature and clinical notes
Guergana K. Savova, Anni Coden, Igor L. Sominsky, Rie Johnson, Philip V. Ogren, Piet C. de Groen, Christopher G. Chute
J. Biomed. Informatics4
2008 Graph-Based Semi-Supervised Learning and Spectral Kernel Design
abstract
In this paper, we consider a framework for semi-supervised learning using spectral decomposition-based unsupervised kernel design. We relate this approach to previously proposed semi-supervised learning methods on graphs. We examine various theoretical properties of such methods. In particular, we present learning bounds and derive optimal kernel representation by minimizing the bound. Based on the theoretical analysis, we are able to demonstrate why spectral kernel design based methods can improve the predictive performance. Empirical examples are included to illustrate the main consequences of our analysis.
Rie Johnson, Tong Zhang 0001
IEEE Trans. Inf. Theory1
2007 On the Effectiveness of Laplacian Normalization for Graph Semi-supervised Learning
Rie Johnson, Tong Zhang 0001
J. Mach. Learn. Res.1