Yanxin Xi

dblp:262/4104 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0003-4715-2186ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 UrbanLLaVA: A Multi-Modal Large Language Model for Urban Intelligence
Jie Feng 0002, Tianhui Liu, Yanxin Xi, Yong Li 0008
ICCV4
2025 Satellites Reveal Mobility: A Commuting Origin-destination Flow Generator for Global Cities
abstract
Commuting Origin-destination (OD) flows, capturing daily population mobility of citizens, are vital for sustainable development across cities around the world. However, it is challenging to obtain the data due to the high cost of travel surveys and privacy concerns. Surprisingly, we find that satellite imagery, publicly available across the globe, contains rich urban semantic signals to support high-quality OD flow generation, with over 98\% expressiveness of traditional multisource hard-to-collect urban sociodemographic, economics, land use, and point of interest data. This inspires us to design a novel data generator, GlODGen (Global-scale OriginDestination Flow Generator), which can generate OD flow data for any cities of interest around the world. Specifically, GlODGen first leverages Vision-Language Geo-Foundation Models to extract urban semantic signals related to human mobility from satellite imagery. These features are then combined with population data to form region-level representations, which are used to generate OD flows via graph diffusion models. Extensive experiments on 4 continents and 6 representative cities show that GlODGen has great generalizability across diverse urban environments on different continents and can generate OD flow data for global cities highly consistent with real-world mobility data. We implement GlODGen as an automated tool, seamlessly integrating data acquisition and curation, urban semantic feature extraction, and OD flow generation together. It has been released at https://github.com/tsinghua-fib-lab/generate-od-pubtools.
Can Rong, Xin Zhang 0106, Yanxin Xi, Hongjie Sui, Jingtao Ding, Yong Li 0008
NeurIPS3
2024 GUI: A Comprehensive Dataset of Global Urban Infrastructure Based on Geospatial Visual Foundation Models
abstract
The substantial social and financial costs of infrastructure identification impede in-depth analyses of sustainable urban design, especially in developing countries. In this paper, we present a novel framework with interactive web visualization based on geospatial visual foundation models. Leveraging this framework, we examine the urban infrastructure information in 1,178 cities worldwide, covering 93, 088 km2 areas. Cross-validation reveals that the overall accuracy of identified infrastructure achieves 67.0%. It sheds light on the sustainable development of cities and exposes the stark inequity in urban infrastructure provision for vulnerable populations. The identified urban infrastructure dataset of this study are available at https://github.com/tsinghua-fib-lab/GUI, and the interactive web application is at https://tinyurl.com/yz7xbfy3.
Zhenyu Han, Xin Zhang 0106, Yanxin Xi, Tong Xia, Yong Li 0008
SIGSPATIAL/GIS3
2024 From Pixels to Progress: Generating Road Network from Satellite Imagery for Socioeconomic Insights in Impoverished Areas
Yanxin Xi, Yu Liu 0016, Sasu Tarkoma, Pan Hui 0001, Yong Li 0008
IJCAI1
2023 Devil in the Landscapes: Inferring Epidemic Exposure Risks from Street View Imagery
abstract
Built environment supports all the daily activities and shapes our health. Leveraging informative street view imagery, previous research has established the profound correlation between the built environment and chronic, non-communicable diseases; however, predicting the exposure risk of infectious diseases remains largely unexplored. The person-to-person contacts and interactions contribute to the complexity of infectious disease, which is inherently different from non-communicable diseases. Besides, the complex relationships between street view imagery and epidemic exposure also hinder accurate predictions. To address these problems, we construct a regional mobility graph informed by the gravity model, based on which we propose a transmission-aware graph neural network (GNN) to capture disease transmission patterns arising from human mobility. Experiments show that the proposed model significantly outperforms baseline models by 8.54% in weighted F1, shedding light on a low-cost, scalable approach to assess epidemic exposure risks from street view imagery.
Zhenyu Han, Yanxin Xi, Tong Xia, Yu Liu 0016, Yong Li 0008
SIGSPATIAL/GIS2
2023 Knowledge-infused Contrastive Learning for Urban Imagery-based Socioeconomic Prediction
abstract
Monitoring sustainable development goals requires accurate and timely socioeconomic statistics, while ubiquitous and frequently-updated urban imagery in web like satellite/street view images has emerged as an important source for socioeconomic prediction. Especially, recent studies turn to self-supervised contrastive learning with manually designed similarity metrics for urban imagery representation learning and further socioeconomic prediction, which however suffers from effectiveness and robustness issues. To address such issues, in this paper, we propose a Knowledge-infused Contrastive Learning (KnowCL) model for urban imagery-based socioeconomic prediction. Specifically, we firstly introduce knowledge graph (KG) to effectively model the urban knowledge in spatiality, mobility, etc., and then build neural network based encoders to learn representations of an urban image in associated semantic and visual spaces, respectively. Finally, we design a cross-modality based contrastive learning framework with a novel image-KG contrastive loss, which maximizes the mutual information between semantic and visual representations for knowledge infusion. Extensive experiments of applying the learnt visual representations for socioeconomic prediction on three datasets demonstrate the superior performance of KnowCL with over 30% improvements on R2 compared with baselines. Especially, our proposed KnowCL model can apply to both satellite and street imagery with both effectiveness and transferability achieved, which provides insights into urban imagery-based socioeconomic prediction.
Yu Liu 0016, Xin Zhang 0106, Jingtao Ding, Yanxin Xi, Yong Li 0008
WWW4
2023 Learning Representations of Satellite Imagery by Leveraging Point-of-Interests
abstract
Satellite imagery depicts the Earth’s surface remotely and provides comprehensive information for many applications, such as land use monitoring and urban planning. Existing studies on unsupervised representation learning for satellite images only take into account the images’ geographic information, ignoring human activity factors. To bridge this gap, we propose using the Point-of-Interest (POI) data to capture human factors and designing a contrastive learning-based framework to consolidate the representation of satellite imagery with POI information. Besides, we introduce a season-invariant representation learning model on satellite imagery, considering that human factors are mostly unchanging with respect to seasons. An attention model is designed at last to merge the representations from the geographic, seasonal, and POI perspectives adaptively. On the basis of real-world datasets collected from Beijing, 1 we evaluate our method for predicting socioeconomic indicators. The results show that the representation containing POI information outperforms the geographic representation in estimating commercial activity-related indicators. Our proposed attentional framework can estimate the socioeconomic indicators with R 2 of 0.874 and outperforms the baseline methods. Furthermore, we explore the differences in the representations of satellite images with varying socioeconomic statuses. Finally, we investigate the impact of geographic and POI perspective information in the representation learning process, as well as the effect of satellite imagery on various spatial resolutions.
Tong Li 0013, Yanxin Xi, Huandong Wang, Yong Li 0008, Sasu Tarkoma, Pan Hui 0001
ACM Trans. Intell. Syst. Technol.2
2022 Predicting Multi-level Socioeconomic Indicators from Structural Urban Imagery
abstract
Understanding economic development and designing government policies requires accurate and timely measurements of socioeconomic activities. In this paper, we show how to leverage city structural information and urban imagery like satellite images and street view images to accurately predict multi-level socioeconomic indicators. Our framework consists of four steps. First, we extract structural information from cities by transforming real-world street networks into city graphs (GeoStruct). Second, we design a contrastive learning-based model to refine urban image features by looking at geographic similarity between images, with images that are geographically close together having similar features (GeoCLR). Third, we propose using street segments as containers to adaptively fuse the features of multi-view urban images, including satellite images and street view images (GeoFuse). Finally, given the city graph with a street segment as a node and a neighborhood area as a subgraph, we jointly model street- and neighborhood-level socioeconomic indicator predictions as node and subgraph classification tasks. The novelty of our method is that we introduce city structure to organize multi-view urban images and model the relationships between socioeconomic indicators at different levels. We evaluate our framework on the basis of real-world datasets collected in multiple cities. Our proposed framework improves performance by over 10% when compared to state-of-the-art baselines in terms of prediction accuracy and recall.
Tong Li 0013, Shiduo Xin, Yanxin Xi, Sasu Tarkoma, Pan Hui 0001, Yong Li 0008
CIKM3
2022 Beyond the First Law of Geography: Learning Representations of Satellite Imagery by Leveraging Point-of-Interests
abstract
Satellite imagery depicts the earth’s surface remotely and provides comprehensive information for many applications, such as land use monitoring and urban planning. Existing studies on unsupervised representation learning for satellite images only take into account the images’ geographic information, ignoring human activity factors. To bridge this gap, we propose using Point-of-Interest (POI) data to capture human factors and design a contrastive learning-based framework to consolidate the representation of satellite imagery with POI information. Also, we design an attention model that merges the representations from the geographic and POI perspectives adaptively. On the basis of real-world datasets collected from Beijing, we evaluate our method for predicting socioeconomic indicators. The results show that the representation containing POI information outperforms the geographic representation in estimating commercial activity-related indicators. Our proposed framework can estimate the socioeconomic indicators with an R2 of 0.874 and outperforms the baseline methods.
Yanxin Xi, Tong Li 0013, Huandong Wang, Yong Li 0008, Sasu Tarkoma, Pan Hui 0001
WWW1
2022 Multitarget Detection Algorithms for Multitemporal Remote Sensing Data
abstract
Target detection is always an important topic in the field of hyper/multispectral remote sensing image processing. At present, target detection algorithms in this field are generally limited to processing single-temporal remote sensing data, and they cannot obtain satisfactory results when the spectra of target and background are similar to each other. Recently, a target detection algorithm called filter tensor analysis (FTA), which is specially designed for multitemporal remote sensing data, has been reported and has achieved better detection results in many cases than the traditional single-temporal methods. However, FTA can only extract one target of interest at a time, and it cannot work when there are multiple targets of interest in the image. Therefore, considering that the matrix form of the FTA method is similar to that of the constrained energy minimization (CEM) model, it naturally comes to us that we can combine the tensor filter in FTA and the multiple target constraints to detect multiple targets by fully exploiting the time-series information in multitemporal data. To be specific, through: 1) adding the “output to one” constraints to the multiple targets in FTA; 2) applying linear/nonlinear function to the outputs of FTA for the multiple targets; and 3) modifying the autocorrelation matrix in FTA, four multitarget detection algorithms for multitemporal remote sensing data are proposed in this article. Experiments with simulation data and real data both show the effectiveness and superiority of the proposed methods.
Yanxin Xi, Luyan Ji, Weitun Yang, Xiurui Geng, Yongchao Zhao
IEEE Trans. Geosci. Remote. Sens.1
2021 FastVGBS: A Fast Version of the Volume-Gradient-Based Band Selection Method for Hyperspectral Imagery
abstract
Recently, the volume-gradient-based band selection (VGBS) method has attracted more and more attention in the field of band selection. It is a ranking-based unsupervised algorithm which applies the sequential backward selection strategy to successively remove the most abundant band. The key finding of VGBS is that the band redundancy corresponds to the volume gradient matrix with respect to hyperspectral images. However, we have found that VGBS requires to update the gradient matrix after each band removal, which includes the calculation of the matrix inverse, determinant, and multiplication, and thus is time-consuming when the number of bands is large. In this letter, we first find that the norm of the row of the gradient matrix has a one-to-one correspondence to the diagonal element of the covariance matrix of the image. Further, we develop a recursive formula to calculate the inverse of the covariance matrix. The experimental results show the effectiveness of the method, i.e., we can reduce the computational complexity of VGBS with an order of magnitude.
Luyan Ji, Liangliang Zhu, Lei Wang 0112, Yanxin Xi, Kai Yu 0006, Xiurui Geng
IEEE Geosci. Remote. Sens. Lett.4