VLDB 2026 Research / reviewers in the wild / expert
Zhili Li
dblp:81/10616
· DBLP profile ↗
21ranked-venue papers
5as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EcoDiffusion: Uncertainty-Aware Emulation of Ecosystem Processes with Conditional Diffusion for Long Sequences with Single-Step InitializationabstractTerrestrial ecosystems constitute a major component of the global carbon sink and play a critical role in regulating the global carbon cycle. Although process-based models such as the Ecosystem Demography (ED) model are widely used to simulate these dynamics and widely adopted in research and applications, they remain computationally intensive and are not well suited for large-scale (e.g., global) projections at high spatial and temporal resolution, or under wide-range of future scenarios. AI-based emulators of process-based physical models have emerged as promising ways to accelerate the computation. However, there are several challenges in developing emulators for ecosystem processes, including error accumulation over long sequences, single-step initial conditions, and high-dimensional environmental conditions. Existing works often rely on time-series patterns in look-back windows, which are not well-suited for the problem with single-step initial conditions. Moreover, they often do not consider uncertainty, making it hard to know when the approximations are highly confident and when the results may need to be updated, e.g., by the process-based models. To address these limitations, we introduce EcoDiffusion, a conditional diffusion framework tailored for ecosystem dynamics emulation. We evaluated EcoDiffusion at locations distributed worldwide under different scenarios and showed that it demonstrated significant improvements over existing models. Xiaowei Jia, Gengchen Mai, George C. Hurtt, Quan Shen, Zhili Li, Yiqun Xie |
AAAI | 8 |
| 2025 | LLM×MapReduce: Simplified Long-Sequence Processing using Large Language ModelsabstractZihan Zhou, Chong Li, Xinyi Chen, Shuo Wang, Yu Chao, Zhili Li, Haoyu Wang, Qi Shi, Zhixing Tan, Xu Han, Xiaodong Shi, Zhiyuan Liu, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shuo Wang 0013, Yu Chao, Zhili Li, Qi Shi 0002, Zhixing Tan, Xu Han 0007, Xiaodong Shi, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 6 |
| 2025 | IsoSim: A Long-term Benchmark Dataset for Water Isotope Emulation in Global Climate ModelsabstractIsotopic ratios of hydrogen and oxygen in water serve as powerful tracers of the Earth's hydrological cycle, offering insights into the origins of water vapor, large-scale atmospheric circulation, and moisture transport dynamics. However, integrating water isotopes into fully coupled global climate models (GCMs) is both scientifically and technically challenging due to the complex interactions between water isotopes and the atmosphere, hydrosphere, and cryosphere, as well as the extensive modifications to model physics and dynamics. As a result, most GCMs lack support for isotopes and even the few existing isotope-enabled GCMs still remain highly expensive to run, significantly limiting their usability. Machine learning (ML) offers promising opportunities to emulate the complex process as powerful mathematical approximators. The water isotope fields from the emulators bring potential for applications in isotope-unenabled GCMs. However, the absence of a publicly available ML-ready dataset has hindered the development of robust ML-based emulators. To address this gap, we introduce IsoSim, the first ML-ready benchmark dataset designed to facilitate the development of ML emulators for water isotopes in GCMs. This dataset includes global climate variables and water isotope fields across three spatial dimensions (latitude, longitude, and height) from isotope-enabled GCM simulations, spanning 500 years at a monthly resolution. We also include different climatic scenarios and a diverse set of learning-based emulators to carry out extensive evaluations and build the benchmarks. The dataset and results serve as reference points to compare machine learning models' ability in approximating complex physical relationships. Zhili Li, Xiaowei Jia, Yiqun Xie |
SIGSPATIAL/GIS | 1 |
| 2025 | Scenario-Based Evaluation of Probabilistic Time Series Forecasting for Solar EnergyabstractProbabilistic time-series forecasting plays a vital role in decisionmaking under uncertainty, especially in applications like solar energy, where forecast reliability directly impacts energy planning and grid stability. While recent models have improved in generating predictive distributions rather than single-point estimates, existing evaluations often focus on average performance and overlook how model quality varies across different real-world scenarios. In solar energy monitoring, for example, the difficulty of forecasting can change significantly due to atmospheric variability, sensor types, and climate conditions. This work addresses the need for scenario-aware evaluation of probabilistic models by benchmarking state-of-the-art forecasting methods using SolarCube-a large-scale solar radiation dataset spanning diverse regions, cloud regimes, and environmental conditions. We define structured "easy" and "hard" cases across four scenarios and examine how different probabilistic model families (e.g., diffusion, VAE, flow-based) capture uncertainty under these conditions. Our goal is to move beyond overall metrics and reveal how model reliability changes across scenarios that are critical for downstream applications. Yiqun Xie, Xiaowei Jia, Gengchen Mai, Sophia Hou, Zhili Li |
SIGSPATIAL/GIS | 7 |
| 2025 | Coincident Data Discovery Engine: A Portal for Global-Scale Cross-Platform Satellite Data SearchabstractCoincident satellite data refer to remote sensing observations from different platforms that capture the same geographic location within a short temporal window. Such data enable multi-modal and multi-view analysis—particularly for dynamic systems like the Arctic, where images captured just hours apart can reflect vastly different conditions (e.g., moving sea ice), complicating data integration. However, acquiring coincident data from different satellite platforms remains labor-intensive and computationally demanding. Finding coincident data is challenging due to separate query systems, inconsistent file formats, and the lack of built-in tools to compute cross-platform time differences. Researchers often settle for loosely aligned observations, compromising temporal alignment precision and analysis quality. To address this challenge, we introduce the Coincident Data Discovery Engine (CoDD), a platform designed to facilitate the discovery and access to spatially and temporally coincident remote sensing data across multiple platforms. CoDD provides data from seven widely used satellite missions spanning optical, SAR, and LiDAR modalities, with a maximum temporal gap of 72 hours and temporal granularity down to one second. The platform features an intuitive user interface that enables efficient querying, visualization, and download of coincident data. CoDD is already supporting diverse ongoing research such as Arctic change monitoring, satellite product validation, and ground truth generation for super-resolution models. Yiqun Xie, Leo Du, Jia Yu 0001, Kyle Duncan, Sinéad Louise Farrell, Zhili Li, Kangyang Chai |
SIGSPATIAL/GIS | 7 |
| 2025 | TreeFinder: A US-Scale Benchmark Dataset for Individual Tree Mortality Monitoring Using High-Resolution Aerial ImageryabstractMonitoring individual tree mortality at scale has been found to be crucial for understanding forest loss, ecosystem resilience, carbon fluxes, and climate-induced impacts. However, the fine-granularity monitoring faces major challenges on both the data and methodology sides because: (1) finding isolated individual-level tree deaths requires high-resolution remote sensing images with broad coverage, and (2) compared to regular geo-objects (e.g., buildings), dead trees often exhibit weaker contrast and high variability across tree types, landscapes and ecosystems. Existing datasets on tree mortality primarily rely on moderate-resolution satellite imagery (e.g., 30m resolution), which aims to detect large-patch wipe-outs but is unable to recognize individual-level tree mortality events. Several efforts have explored alternatives via very-high-resolution drone imagery. However, drone images are highly expensive and can only be collected at local scales, which are therefore not suitable for national-scale applications and beyond. To bridge the gaps,we introduce TreeFinder, the first high-resolution remote sensing benchmark dataset designed for individual-level tree mortality mapping across the Contiguous United States (CONUS). Specifically, the dataset uses NAIP imagery at 0.6m resolution that provides wall-to-wall coverage of the entire CONUS. TreeFinder contains images with pixel-level labels generated via extensive manual annotation that covers forested areas in 48 states with over 23,000 hectares. All annotations are rigorously validated using multi-temporal NAIP images and auxiliary vegetation indices from remote sensing imagery. Moreover, TreeFinder includes multiple evaluation scenarios to test the models' ability in generalizing across different geographic regions, climate zones, and forests with different plant function types. Finally, we develop benchmarks using a suite of semantic segmentation models, including both convolutional architectures and more recent foundation models based on vision transformers for general and remote sensing images. Our dataset and code are publicly available on Kaggle and GitHub: https://www.kaggle.com/datasets/zhihaow/tree-finder and https://github.com/zhwang0/treefinder. Cooper Li, George C. Hurtt, Xiaowei Jia, Gengchen Mai, Zhili Li, Yiqun Xie |
NeurIPS | 8 |
| 2025 | CarbonGlobe: A Global-Scale, Multi-Decade Dataset and Benchmark for Carbon Forecasting in Forest EcosystemsabstractForest ecosystems play a critical role in the Earth system as major carbon sinks that are essential for carbon neutralization and climate change mitigation. However, the Earth has undergone significant deforestation and forest degradation, and the remaining forested areas are also facing increasing pressures from socioeconomic factors and climate change, potentially pushing them towards tipping points.Responding to the grand challenge, a theory-based Ecosystem Demography (ED) model has been continuously developed over the past two decades and serves as a key component in major initiatives, including the Global Carbon Budget, NASA Carbon Monitoring System, and US Greenhouse Gas Center. Despite its growing importance in combating climate change and shaping carbon policies, ED's expensive computation significantly limits its ability to estimate carbon dynamics at the global scale with high spatial resolution.Recently, machine learning (ML) models have shown promising potential in approximating theory-based models with interesting success in various domains including weather forecasting, thanks to the open-source benchmark datasets made available.However, there are currently no publicly available ML-ready datasets for global carbon dynamics forecasting in forest ecosystems. The limited data availability hinders the development of corresponding ML emulators. Furthermore, the inputs needed for running ED are highly complex with over a hundred variables from various remote sensing products. To bridge the gap, we develop a new ML-ready benchmark dataset, \textit{CarbonGlobe}, for carbon dynamics forecasting, featuring that: (1) the data has a global-scale coverage at 0.5$^\circ$ resolution; (2) the temporal range spans 40 years; (3) the inputs integrate extensive multi-source data from different sensing products, with calibrated outputs from ED; (4) the data is formatted in ML-ready forms and split into different evaluation scenarios based on climate conditions, etc.; (5) a set of problem-driven metrics is designed to develop benchmarks using various ML models to best align with the needs of downstream applications. Our dataset and code are publicly available on Kaggle and GitHub: https://www.kaggle.com/datasets/zhihaow/carbonglobe and https://github.com/zhwang0/carbon-globe. George C. Hurtt, Xiaowei Jia, Zhili Li, Shuo Xu 0002, Yiqun Xie |
NeurIPS | 7 |
| 2025 | Machine-learning-enabled spatial pattern mining: evaluating the impact of imperfect inputsabstractSpatial pattern mining (SPM) aims to detect geographic locations or areas that present interesting, nontrivial, and potentially useful patterns. Traditional formulations of point-based SPM tasks are mainly based on true observations, which tend to have limited spatial coverage, availability, and timeliness. While machine learning (ML) has the potential to extend the range of usable data, the uncertainty of model-predicted labels presents new challenges for their usability in the SPM context. This paper formulates the task of ML-enabled SPM using predicted labels by ML models. Given the ever-expanding family of spatial patterns, we consider four widely-adopted patterns – hotspots, co-locations, mixture patterns, and spatial outliers – to scope our study to make the discussion concrete. We develop soft-label versions of SPM algorithms that can directly execute on uncertain predictions generated by ML models. Additionally, we evaluate the ML-enabled SPM results for both categorical and real-valued datasets across a spectrum of prediction quality. The results show that certain spatial patterns such as multinomial scan statistic-based mixture patterns and normal-model-based hotspots can more robustly maintain the detection quality at different error levels, while others such as spatial outliers are more sensitive to incorrect predictions. This provides helpful guidance on using learning-based predictions for SPM. Zhili Li, Yiqun Xie, Xiaowei Jia, Gengchen Mai, Weiye Chen |
Int. J. Geogr. Inf. Sci. | 1 |
| 2024 | SimFair: Physics-Guided Fairness-Aware Learning with Simulation ModelsabstractFairness-awareness has emerged as an essential building block for the responsible use of artificial intelligence in real applications. In many cases, inequity in performance is due to the change in distribution over different regions. While techniques have been developed to improve the transferability of fairness, a solution to the problem is not always feasible with no samples from the new regions, which is a bottleneck for pure data-driven attempts. Fortunately, physics-based mechanistic models have been studied for many problems with major social impacts. We propose SimFair, a physics-guided fairness-aware learning framework, which bridges the data limitation by integrating physical-rule-based simulation and inverse modeling into the training design. Using temperature prediction as an example, we demonstrate the effectiveness of the proposed SimFair in fairness preservation. Yiqun Xie, Zhili Li, Xiaowei Jia, Zhe Jiang 0001, Aolin Jia, Shuo Xu 0002 |
AAAI | 3 |
| 2024 | High-Resolution Poverty Mapping with Foundation Models: A Cost-effective Approach from Street Views to Satellite ImagesabstractAlthough standards of living are increasing rapidly worldwide, a considerable segment of the global population continues to live in poverty. Local governments and decision makers urgently need actionable fine-scale poverty maps to know the locations of the low income population for operational resource distribution. However, most existing studies focus on coarse-resolution poverty maps (e.g., county level) and offer limited information to help deliver the resources to the right locations. Moreover, coarse-resolution maps generated by machine learning models are often trained on higher-level economic statistics that have greater availability. However, such labels at the fine-scale remain very scarce, and existing maps are commonly based on household-level visits that are highly expensive and time-consuming, making them only available in a limited number of cities. We develop a cost-effective approach to tackle the challenge. First, we design a multi-view training data construction approach using data from both street views and very-high-resolution satellite images. Next, we integrate different types of foundation models including the general-purpose vision transformer ViT and the segmentation-focused SegFormer for training and map generation in new cities. Via the use of pretrained large models, the goal is to enhance the generalizability with a smaller amount of samples. To validate the approach, we carried out a case study in Ghana with the cities of Accra, Kumasi, and Tamale. The results showed the effectiveness of the cost-effective approach in capturing low-income areas with unique characteristics, and the foundation models also demonstrated enhanced ability in generalization with smaller training data sizes. William Lu, Zhili Li, Yiqun Xie |
IEEE Big Data | 2 |
| 2024 | SolarCube: An Integrative Benchmark Dataset Harnessing Satellite and In-situ Observations for Large-scale Solar Energy ForecastingabstractSolar power is a critical source of renewable energy, offering significant potential to lower greenhouse gas emissions and mitigate climate change. However, the cloud induced-variability of solar radiation reaching the earth’s surface presents a challenge for integrating solar power into the grid (e.g., storage and backup management). The new generation of geostationary satellites such as GOES-16 has become an important data source for large-scale and high temporal frequency solar radiation forecasting. However, no machine-learning-ready dataset has integrated geostationary satellite data with fine-grained solar radiation information to support forecasting model development and benchmarking with consistent metrics. We present SolarCube, a new ML-ready benchmark dataset for solar radiation forecasting. SolarCube covers 19 study areas distributed over multiple continents: North America, South America, Asia, and Oceania. The dataset supports short (i.e., 30 minutes to 6 hours) and long-term (i.e., day-ahead or longer) solar radiation forecasting at both point-level (i.e., specific locations of monitoring stations) and area-level, by processing and integrating data from multiple sources, including geostationary satellite images, physics-derived solar radiation, and ground station observations from different monitoring networks over the globe. We also evaluated a set of forecasting models for point- and image-based time-series data to develop performance benchmarks under different testing scenarios. The dataset is available at https://doi.org/10.5281/zenodo.11498739. A Python library is available to conveniently generate different variations of the dataset based on user needs, along with baseline models at https://github.com/Ruohan-Li/SolarCube. Yiqun Xie, Xiaowei Jia, Dongdong Wang 0001, Yingxue Zhang 0002, Zhili Li |
NeurIPS | 8 |
| 2023 | Point-to-Region Co-learning for Poverty Mapping at High Resolution Using Satellite ImageryabstractDespite improvements in safe water and sanitation services in low-income countries, a substantial proportion of the population in Africa still does not have access to these essential services. Up-to-date fine-scale maps of low-income settlements are urgently needed by authorities to improve service provision. We aim to develop a cost-effective solution to generate fine-scale maps of these vulnerable populations using multi-source public information. The problem is challenging as ground-truth maps are available at only a limited number of cities, and the patterns are heterogeneous across cities. Recent attempts tackling the spatial heterogeneity issue focus on scenarios where true labels partially exist for each input region, which are unavailable for the present problem. We propose a dynamic point-to-region co-learning framework to learn heterogeneity patterns that cannot be reflected by point-level information and generalize deep learners to new areas with no labels. We also propose an attention-based correction layer to remove spurious signatures, and a region-gate to capture both region-invariant and variant patterns. Experiment results on real-world fine-scale data in three cities of Kenya show that the proposed approach can largely improve model performance on various base network architectures. Zhili Li, Yiqun Xie, Xiaowei Jia, Kara Stuart, Caroline Delaire, Serhiy Skakun |
AAAI | 1 |
| 2023 | Auto-CM: Unsupervised Deep Learning for Satellite Imagery Composition and Cloud Masking Using Spatio-Temporal DynamicsabstractCloud masking is both a fundamental and a critical task in the vast majority of Earth observation problems across social sectors, including agriculture, energy, water, etc. The sheer volume of satellite imagery to be processed has fast-climbed to a scale (e.g., >10 PBs/year) that is prohibitive for manual processing. Meanwhile, generating reliable cloud masks and image composite is increasingly challenging due to the continued distribution-shifts in the imagery collected by existing sensors and the ever-growing variety of sensors and platforms. Moreover, labeled samples are scarce and geographically limited compared to the needs in real large-scale applications. In related work, traditional remote sensing methods are often physics-based and rely on special spectral signatures from multi- or hyper-spectral bands, which are often not available in data collected by many -- and especially more recent -- high-resolution platforms. Machine learning and deep learning based methods, on the other hand, often require large volumes of up-to-date training data to be reliable and generalizable over space. We propose an autonomous image composition and masking (Auto-CM) framework to learn to solve the fundamental tasks in a label-free manner, by leveraging different dynamics of events in both geographic domains and time-series. Our experiments show that Auto-CM outperforms existing methods on a wide-range of data with different satellite platforms, geographic regions and bands. Yiqun Xie, Zhili Li, Han Bao 0003, Xiaowei Jia, Dongkuan Xu, Xun Zhou 0001, Serhiy Skakun |
AAAI | 2 |
| 2023 | Confidence-based Self-Corrective Learning: An Application in Height Estimation Using Satellite LiDAR and ImageryabstractWidespread, and rapid, environmental transformation is underway on Earth driven by human activities. Climate shifts such as global warming have led to massive and alarming loss of ice and snow in the high-latitude regions including the Arctic, causing many natural disasters due to sea-level rise, etc. Mitigating the impacts of climate change has also become a United Nations' Sustainable Development Goal for 2030. The recent launch of the ICESat-2 satellites target on heights in the polar regions. However, the observations are only available along very narrow scan lines, leaving large no-data gaps in-between. We aim to fill the gaps by combining the height observations with high-resolution satellite imagery that have large footprints (spatial coverage). The data expansion is a challenging task as the height data are often constrained on one or a few lines per image in real applications, and the images are highly noisy for height estimation. Related work on image-based height prediction and interpolation relies on specific types of images or does not consider the highly-localized height distribution. We propose a spatial self-corrective learning framework, which explicitly uses confidence-based pseudo-interpolation, recurrent self-refinement, and truth-based correction with a regression layer to address the challenges. We carry out experiments on different landscapes in the high-latitude regions and the proposed method shows stable improvements compared to the baseline methods. Zhili Li, Yiqun Xie, Xiaowei Jia |
IJCAI | 1 |
| 2023 | Current Progress and Challenges in Large-Scale 3D Mitochondria Instance SegmentationabstractIn this paper, we present the results of the MitoEM challenge on mitochondria 3D instance segmentation from electron microscopy images, organized in conjunction with the IEEE-ISBI 2021 conference. Our benchmark dataset consists of two large-scale 3D volumes, one from human and one from rat cortex tissue, which are 1,986 times larger than previously used datasets. At the time of paper submission, 257 participants had registered for the challenge, 14 teams had submitted their results, and six teams participated in the challenge workshop. Here, we present eight top-performing approaches from the challenge participants, along with our own baseline strategies. Posterior to the challenge, annotation errors in the ground truth were corrected without altering the final ranking. Additionally, we present a retrospective evaluation of the scoring system which revealed that: 1) challenge metric was permissive with the false positive predictions; and 2) size-based grouping of instances did not correctly categorize mitochondria of interest. Thus, we propose a new scoring system that better reflects the correctness of the segmentation results. Although several of the top methods are compared favorably to our own baselines, substantial errors remain unsolved for mitochondria with challenging morphologies. Thus, the challenge remains open for submission and automatic evaluation, with all volumes available for download. Daniel Franco-Barranco, Zudi Lin, Won-Dong Jang, Xueying Wang 0002, Qijia Shen, Yutian Fan, Mingxing Li 0003, Chang Chen 0004, Zhiwei Xiong, Rui Xin 0003, Huai Chen, Zhili Li, Jie Zhao 0020, Xuejin Chen, Constantin Pape, Ryan Conrad, Luke Nightingale, Joost de Folter, Martin L. Jones, Dorsa Ziaei, Stephan Huschauer, Ignacio Arganda-Carreras, Hanspeter Pfister, Donglai Wei 0001 |
IEEE Trans. Medical Imaging | 14 |
| 2022 | Deep semantic segmentation for building detection using knowledge-informed features from LiDAR point cloudsabstractAirborne LiDAR point clouds record three-dimensional structures of ground surfaces with high precision, and have been widely used to identify geospatial objects, facilitating the understanding of the distribution and changing dynamics of the environment. Detection can be complicated by the complex structures of ground objects and noises in LiDAR point clouds. Related work has explored the use of deep learning techniques such as YOLO in detecting geospatial objects (e.g., building footprints) on both optical imagery and LiDAR point clouds. However, deep networks are data hungry and there are often limited labeled samples available for many geospatial object mapping tasks, making it difficult for the models to generalize to unseen test regions. This paper describes the framework used in the 11th SIGSPATIAL Cup Competition (GIS CUP 2022), which received the top-3 performance. Our framework incorporates domain knowledge to reduce the difficulty of learning and the model's reliance on large training sets. Specifically, we present knowledge-informed feature generation and filtering based on morphological characteristics to improve the generalizability of learned features. Then, we use a deep segmentation backbone (U-Net) with training- and test-time augmentation to generate preliminary candidates for building footprints. Finally, we utilize domain rules (e.g., geometric properties) to regularize and filter the detections to create the final map of building footprints. Experiment results show that the strategies can effectively improve detection results in different landscapes. Weiye Chen, Zhili Li, Yiqun Xie, Xiaowei Jia, Anlin Li |
SIGSPATIAL/GIS | 3 |
| 2020 | RADC-Net: A residual attention based convolution network for aerial scene classification
Qi Bi, Kun Qin, Han Zhang 0052, Zhili Li |
Neurocomputing | 4 |
| 2020 | APDC-Net: Attention Pooling-Based Convolutional Network for Aerial Scene ClassificationabstractDeep learning methods have boosted the performance of a series of visual tasks. However, the aerial image scene classification remains challenging. The object distribution and spatial arrangement in aerial scenes are often more complicated than in natural image scenes. Possible solutions include highlighting local semantics relevant to the scene label and preserving more discriminative features. To tackle this challenge, in this letter, we propose an attention pooling-based dense connected convolutional network (APDC-Net) for aerial scene classification. First, it uses a simplified dense connection structure as the backbone to preserve features from different levels. Then, we propose a trainable pooling to down-sample the feature maps and to enhance the local semantic representation capability. Finally, we introduce a multi-level supervision strategy, so that features from different levels are all allowed to supervise the training process directly. Exhaustive experiments on three aerial scene classification benchmarks demonstrate that our proposed APDC-Net outperforms other state-of-the-art methods with much fewer parameters and validate the effectiveness of our attention-based pooling and multi-level supervision strategy. Qi Bi, Kun Qin, Han Zhang 0052, Jiafen Xie, Zhili Li |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2020 | Deep Multiple Instance Convolutional Neural Networks for Learning Robust Scene RepresentationsabstractThe accuracy and efficiency of scene classification have immensely improved with the extensive application of deep convolutional neural networks (CNNs). However, standard CNNs classify images mostly based on the global features from the last fully connected layer, which may cause the negligence of discriminative local information and the sensitivity to various spatial transformations. In this article, we consider the problem of scene classification from the perspective of multiple instance learning (MIL) and propose an end-to-end multiple instance CNN (MI-CNN) for learning more robust scene representations. In MI-CNN, a scene is represented as a bag of local patches (instances). An instance-level classifier is trained to obtain the label of each patch in an MIL fashion, which makes the classifier more sensitive to the discriminative local patches. The patch labels are then aggregated into an image label by an MIL pooling layer, which is invariant to the order of local patches and helps construct more robust representations. We present extensive experiments on UC Merced Land use (UCM), Aerial Image data set (AID), and NWPU-RESISC (NWPU) data sets. Experimental results show that the proposed method achieves 1.17%, 1.70%, and 3.61% accuracy improvements with 90% parameter reduction compared with the standard CNNs. Zhili Li, Jiafen Xie, Qi Bi, Kun Qin |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | A Multiple-Instance Densely-Connected ConvNet for Aerial Scene ClassificationabstractIn contrast with nature scenes, aerial scenes are often composed of many objects crowdedly distributed on the surface in bird's view, the description of which usually demands more discriminative features as well as local semantics. However, when applied to scene classification, most of the existing convolution neural networks (ConvNets) tend to depict global semantics of images, and the loss of low- and mid-level features can hardly be avoided, especially when the model goes deeper. To tackle these challenges, in this paper, we propose a multiple-instance densely-connected ConvNet (MIDC-Net) for aerial scene classification. It regards aerial scene classification as a multiple-instance learning problem so that local semantics can be further investigated. Our classification model consists of an instance-level classifier, a multiple instance pooling and followed by a bag-level classification layer. In the instance-level classifier, we propose a simplified dense connection structure to effectively preserve features from different levels. The extracted convolution features are further converted into instance feature vectors. Then, we propose a trainable attention-based multiple instance pooling. It highlights the local semantics relevant to the scene label and outputs the bag-level probability directly. Finally, with our bag-level classification layer, this multiple instance learning framework is under the direct supervision of bag labels. Experiments on three widely-utilized aerial scene benchmarks demonstrate that our proposed method outperforms many state-of-the-art methods by a large margin with much fewer parameters. Qi Bi, Kun Qin, Zhili Li, Han Zhang 0052, Gui-Song Xia |
IEEE Trans. Image Process. | 3 |
| 2019 | Multiple Instance Dense Connected Convolution Neural Network for Aerial Image Scene ClassificationabstractWith the development of deep learning, many state-of-the-art natural image scene classification methods have demonstrated impressive performance. While the current convolution neural network tends to extract global features and global semantic information in a scene, the geo-spatial objects can be located at anywhere in an aerial image scene and their spatial arrangement tends to be more complicated. One possible solution is to preserve more local semantic information and enhance feature propagation. In this paper, an end to end multiple instance dense connected convolution neural network (MIDCCNN) is proposed for aerial image scene classification. First, a 23 layer dense connected convolution neural network (DCCNN) is built and served as a backbone to extract convolution features. It is capable of preserving middle and low level convolution features. Then, an attention based multiple instance pooling is proposed to highlight the local semantics in an aerial image scene. Finally, we minimize the loss between the bag-level predictions and the ground truth labels so that the whole framework can be trained directly. Experiments on three aerial image datasets demonstrate that our proposed methods can outperform current baselines by a large margin. Qi Bi, Kun Qin, Zhili Li, Han Zhang 0052 |
ICIP | 3 |