Neal Jean

dblp:169/9869 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
0since 2021 · last 2019
0000-0002-2616-3348ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Representation and self-supervised learning · 34% Probabilistic and Bayesian machine learning · 26% Learning paradigms · 26%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational social science and digital humanities · 100%

Topics — the 10 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning
spatial representation learning
0.412019
Tile2Vec: Unsupervised Representation Learning for Spatially Distributed Data · AAAI 2019
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning
0.412019
Tile2Vec: Unsupervised Representation Learning for Spatially Distributed Data · AAAI 2019
Computational social science and digital humanities
socioeconomic indicator prediction
0.412019
Predicting Economic Development using Geolocated Wikipedia Articles · KDD 2019
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process › kernel design
deep kernel learning
0.312018
Semi-supervised Deep Kernel Learning: Regression with Unlabeled Data by Minimizing Predictive Variance · NeurIPS 2018
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.312018
Semi-supervised Deep Kernel Learning: Regression with Unlabeled Data by Minimizing Predictive Variance · NeurIPS 2018
Machine learning › Learning paradigms
semi-supervised learning
0.312018
Semi-supervised Deep Kernel Learning: Regression with Unlabeled Data by Minimizing Predictive Variance · NeurIPS 2018
Machine learning › Learning paradigms › semi-supervised learning
semi-supervised regression
0.312018
Semi-supervised Deep Kernel Learning: Regression with Unlabeled Data by Minimizing Predictive Variance · NeurIPS 2018
Computer vision › 3D vision › remote sensing
remote sensing image analysis
0.212016
Transfer Learning from Deep Features for Remote Sensing and Poverty Mapping · AAAI 2016
Machine learning › Representation and self-supervised learning
distributional hypothesis
0.112019
Tile2Vec: Unsupervised Representation Learning for Spatially Distributed Data · AAAI 2019
Computational social science and digital humanities › socioeconomic indicator prediction
poverty mapping
0.112016
Transfer Learning from Deep Features for Remote Sensing and Poverty Mapping · AAAI 2016

Methods — techniques the papers use, named apart from their topics

satellite imagery analysis · 0.8natural language processing · 0.8fully convolutional network · 0.5convolutional neural network · 0.5predictive variance minimization · 0.3posterior regularization · 0.3gaussian process · 0.3
YearPublicationVenuePosition
2019 Tile2Vec: Unsupervised Representation Learning for Spatially Distributed Data
abstract
Geospatial analysis lacks methods like the word vector representations and pre-trained networks that significantly boost performance across a wide range of natural language and computer vision tasks. To fill this gap, we introduce Tile2Vec, an unsupervised representation learning algorithm that extends the distributional hypothesis from natural language — words appearing in similar contexts tend to have similar meanings — to spatially distributed data. We demonstrate empirically that Tile2Vec learns semantically meaningful representations for both image and non-image datasets. Our learned representations significantly improve performance in downstream classification tasks and, similarly to word vectors, allow visual analogies to be obtained via simple arithmetic in the latent space.
Neal Jean, Sherrie Wang, Anshul Samar, George Azzari, David B. Lobell, Stefano Ermon
AAAI1
2019 Predicting Economic Development using Geolocated Wikipedia Articles
abstract
Progress on the UN Sustainable Development Goals (SDGs) is hampered by a persistent lack of data regarding key social, environmental, and economic indicators, particularly in developing countries. For example, data on poverty - the first of seventeen SDGs - is both spatially sparse and infrequently collected in Sub-Saharan Africa due to the high cost of surveys. Here we propose a novel method for estimating socioeconomic indicators using open-source, geolocated textual information from Wikipedia articles. We demonstrate that modern NLP techniques can be used to predict community-level asset wealth and education outcomes using nearby geolocated Wikipedia articles. When paired with nightlights satellite imagery, our method outperforms all previously published benchmarks for this prediction task, indicating the potential of Wikipedia to inform both research in the social sciences and future policy decisions.
Evan Sheehan, Chenlin Meng, Matthew Tan, Burak Uzkent, Neal Jean, Marshall Burke, David B. Lobell, Stefano Ermon
KDD5
2018 Semi-supervised Deep Kernel Learning: Regression with Unlabeled Data by Minimizing Predictive Variance
abstract
Large amounts of labeled data are typically required to train deep learning models. For many real-world problems, however, acquiring additional data can be expensive or even impossible. We present semi-supervised deep kernel learning (SSDKL), a semi-supervised regression model based on minimizing predictive variance in the posterior regularization framework. SSDKL combines the hierarchical representation learning of neural networks with the probabilistic modeling capabilities of Gaussian processes. By leveraging unlabeled data, we show improvements on a diverse set of real-world regression tasks over supervised deep kernel learning and semi-supervised methods such as VAT and mean teacher adapted for regression.
Neal Jean, Sang Michael Xie, Stefano Ermon
NeurIPS1
2016 Transfer Learning from Deep Features for Remote Sensing and Poverty Mapping
abstract
The lack of reliable data in developing countries is a major obstacle to sustainable development, food security, and disaster relief. Poverty data, for example, is typically scarce, sparse in coverage, and labor-intensive to obtain. Remote sensing data such as high-resolution satellite imagery, on the other hand, is becoming increasingly available and inexpensive. Unfortunately, such data is highly unstructured and currently no techniques exist to automatically extract useful insights to inform policy decisions and help direct humanitarian efforts. We propose a novel machine learning approach to extract large-scale socioeconomic indicators from high-resolution satellite imagery. The main challenge is that training data is very scarce, making it difficult to apply modern techniques such as Convolutional Neural Networks (CNN). We therefore propose a transfer learning approach where nighttime light intensities are used as a data-rich proxy. We train a fully convolutional CNN model to predict nighttime lights from daytime imagery, simultaneously learning features that are useful for poverty prediction. The model learns filters identifying different terrains and man-made structures, including roads, buildings, and farmlands, without any supervision beyond nighttime lights. We demonstrate that these learned features are highly informative for poverty mapping, even approaching the predictive performance of survey data collected in the field.
Sang Michael Xie, Neal Jean, Marshall Burke, David B. Lobell, Stefano Ermon
AAAI2