VLDB 2026 Research / reviewers in the wild / expert
Neal Jean
dblp:169/9869
· DBLP profile ↗
4ranked-venue papers
2as first author
0since 2021 · last 2019
0000-0002-2616-3348ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Representation and self-supervised learning · 34% Probabilistic and Bayesian machine learning · 26% Learning paradigms · 26% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Computational social science and digital humanities · 100% |
Topics — the 10 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning
spatial representation learning |
0.4 | 1 | 2019 | Tile2Vec: Unsupervised Representation Learning for Spatially Distributed Data · AAAI 2019 |
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning |
0.4 | 1 | 2019 | Tile2Vec: Unsupervised Representation Learning for Spatially Distributed Data · AAAI 2019 |
Computational social science and digital humanities
socioeconomic indicator prediction |
0.4 | 1 | 2019 | Predicting Economic Development using Geolocated Wikipedia Articles · KDD 2019 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process › kernel design
deep kernel learning |
0.3 | 1 | 2018 | Semi-supervised Deep Kernel Learning: Regression with Unlabeled Data by Minimizing Predictive Variance · NeurIPS 2018 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.3 | 1 | 2018 | Semi-supervised Deep Kernel Learning: Regression with Unlabeled Data by Minimizing Predictive Variance · NeurIPS 2018 |
Machine learning › Learning paradigms
semi-supervised learning |
0.3 | 1 | 2018 | Semi-supervised Deep Kernel Learning: Regression with Unlabeled Data by Minimizing Predictive Variance · NeurIPS 2018 |
Machine learning › Learning paradigms › semi-supervised learning
semi-supervised regression |
0.3 | 1 | 2018 | Semi-supervised Deep Kernel Learning: Regression with Unlabeled Data by Minimizing Predictive Variance · NeurIPS 2018 |
Computer vision › 3D vision › remote sensing
remote sensing image analysis |
0.2 | 1 | 2016 | Transfer Learning from Deep Features for Remote Sensing and Poverty Mapping · AAAI 2016 |
Machine learning › Representation and self-supervised learning
distributional hypothesis |
0.1 | 1 | 2019 | Tile2Vec: Unsupervised Representation Learning for Spatially Distributed Data · AAAI 2019 |
Computational social science and digital humanities › socioeconomic indicator prediction
poverty mapping |
0.1 | 1 | 2016 | Transfer Learning from Deep Features for Remote Sensing and Poverty Mapping · AAAI 2016 |
Methods — techniques the papers use, named apart from their topics
satellite imagery analysis · 0.8natural language processing · 0.8fully convolutional network · 0.5convolutional neural network · 0.5predictive variance minimization · 0.3posterior regularization · 0.3gaussian process · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Tile2Vec: Unsupervised Representation Learning for Spatially Distributed DataabstractGeospatial analysis lacks methods like the word vector representations and pre-trained networks that significantly boost performance across a wide range of natural language and computer vision tasks. To fill this gap, we introduce Tile2Vec, an unsupervised representation learning algorithm that extends the distributional hypothesis from natural language — words appearing in similar contexts tend to have similar meanings — to spatially distributed data. We demonstrate empirically that Tile2Vec learns semantically meaningful representations for both image and non-image datasets. Our learned representations significantly improve performance in downstream classification tasks and, similarly to word vectors, allow visual analogies to be obtained via simple arithmetic in the latent space. Neal Jean, Sherrie Wang, Anshul Samar, George Azzari, David B. Lobell, Stefano Ermon |
AAAI | 1 |
| 2019 | Predicting Economic Development using Geolocated Wikipedia ArticlesabstractProgress on the UN Sustainable Development Goals (SDGs) is hampered by a persistent lack of data regarding key social, environmental, and economic indicators, particularly in developing countries. For example, data on poverty - the first of seventeen SDGs - is both spatially sparse and infrequently collected in Sub-Saharan Africa due to the high cost of surveys. Here we propose a novel method for estimating socioeconomic indicators using open-source, geolocated textual information from Wikipedia articles. We demonstrate that modern NLP techniques can be used to predict community-level asset wealth and education outcomes using nearby geolocated Wikipedia articles. When paired with nightlights satellite imagery, our method outperforms all previously published benchmarks for this prediction task, indicating the potential of Wikipedia to inform both research in the social sciences and future policy decisions. Evan Sheehan, Chenlin Meng, Matthew Tan, Burak Uzkent, Neal Jean, Marshall Burke, David B. Lobell, Stefano Ermon |
KDD | 5 |
| 2018 | Semi-supervised Deep Kernel Learning: Regression with Unlabeled Data by Minimizing Predictive VarianceabstractLarge amounts of labeled data are typically required to train deep learning models. For many real-world problems, however, acquiring additional data can be expensive or even impossible. We present semi-supervised deep kernel learning (SSDKL), a semi-supervised regression model based on minimizing predictive variance in the posterior regularization framework. SSDKL combines the hierarchical representation learning of neural networks with the probabilistic modeling capabilities of Gaussian processes. By leveraging unlabeled data, we show improvements on a diverse set of real-world regression tasks over supervised deep kernel learning and semi-supervised methods such as VAT and mean teacher adapted for regression. Neal Jean, Sang Michael Xie, Stefano Ermon |
NeurIPS | 1 |
| 2016 | Transfer Learning from Deep Features for Remote Sensing and Poverty MappingabstractThe lack of reliable data in developing countries is a major obstacle to sustainable development, food security, and disaster relief. Poverty data, for example, is typically scarce, sparse in coverage, and labor-intensive to obtain. Remote sensing data such as high-resolution satellite imagery, on the other hand, is becoming increasingly available and inexpensive. Unfortunately, such data is highly unstructured and currently no techniques exist to automatically extract useful insights to inform policy decisions and help direct humanitarian efforts. We propose a novel machine learning approach to extract large-scale socioeconomic indicators from high-resolution satellite imagery. The main challenge is that training data is very scarce, making it difficult to apply modern techniques such as Convolutional Neural Networks (CNN). We therefore propose a transfer learning approach where nighttime light intensities are used as a data-rich proxy. We train a fully convolutional CNN model to predict nighttime lights from daytime imagery, simultaneously learning features that are useful for poverty prediction. The model learns filters identifying different terrains and man-made structures, including roads, buildings, and farmlands, without any supervision beyond nighttime lights. We demonstrate that these learned features are highly informative for poverty mapping, even approaching the predictive performance of survey data collected in the field. Sang Michael Xie, Neal Jean, Marshall Burke, David B. Lobell, Stefano Ermon |
AAAI | 2 |