EDBT 2026 Demo / reviewers in the wild / expert
Xin Zhang 0106
dblp:76/1584-106
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0002-2506-7370ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CityBench: Evaluating the Capabilities of Large Language Models for Urban TasksabstractAs large language models (LLMs) continue to advance and gain widespread use, establishing systematic and reliable evaluation methodologies for LLMs and vision-language models (VLMs) has become essential to ensure their real-world effectiveness and reliability. There have been some early explorations about the usability of LLMs for limited urban tasks, but a systematic and scalable evaluation benchmark is still lacking. The challenge in constructing a systematic evaluation benchmark for urban research lies in the diversity of urban data, the complexity of application scenarios and the highly dynamic nature of the urban environment. In this paper, we design CityBench, an interactive simulator based evaluation platform, as the first systematic benchmark for evaluating the capabilities of LLMs for diverse tasks in urban research. First, we build CityData to integrate the diverse urban data and CitySimu to simulate fine-grained urban dynamics. Based on CityData and CitySimu, we design 8 representative urban tasks in 2 categories of perception-understanding and decision-making as the CityBench. With extensive results from 30 well-known LLMs and VLMs in 13 cities around the world, we find that advanced LLMs and VLMs can achieve competitive performance in diverse urban tasks requiring commonsense and semantic understanding abilities, e.g., understanding the human dynamics and semantic inference of urban images. Meanwhile, they fail to solve the challenging urban tasks requiring professional knowledge and high-level numerical abilities, e.g., geospatial prediction and traffic control task. These findings provide critical insights for the effective utilization and further development of LLMs to advance urban-related tasks and research in the future. Jie Feng 0002, Jun Zhang 0087, Tianhui Liu, Xin Zhang 0106, Tianjian Ouyang, Junbo Yan, Yuwei Du, Yong Li 0008 |
KDD (2) | 4 |
| 2025 | Satellites Reveal Mobility: A Commuting Origin-destination Flow Generator for Global CitiesabstractCommuting Origin-destination (OD) flows, capturing daily population mobility of citizens, are vital for sustainable development across cities around the world. However, it is challenging to obtain the data due to the high cost of travel surveys and privacy concerns. Surprisingly, we find that satellite imagery, publicly available across the globe, contains rich urban semantic signals to support high-quality OD flow generation, with over 98\% expressiveness of traditional multisource hard-to-collect urban sociodemographic, economics, land use, and point of interest data. This inspires us to design a novel data generator, GlODGen (Global-scale OriginDestination Flow Generator), which can generate OD flow data for any cities of interest around the world. Specifically, GlODGen first leverages Vision-Language Geo-Foundation Models to extract urban semantic signals related to human mobility from satellite imagery. These features are then combined with population data to form region-level representations, which are used to generate OD flows via graph diffusion models. Extensive experiments on 4 continents and 6 representative cities show that GlODGen has great generalizability across diverse urban environments on different continents and can generate OD flow data for global cities highly consistent with real-world mobility data. We implement GlODGen as an automated tool, seamlessly integrating data acquisition and curation, urban semantic feature extraction, and OD flow generation together. It has been released at https://github.com/tsinghua-fib-lab/generate-od-pubtools. Can Rong, Xin Zhang 0106, Yanxin Xi, Hongjie Sui, Jingtao Ding, Yong Li 0008 |
NeurIPS | 2 |
| 2025 | Perceiving Urban Inequality from Imagery Using Visual Language Models with Chain-of-Thought ReasoningabstractThe rapid pace of urbanization has led to unequal benefits for residents, creating significant inequality issues and discussions around Sustainable Development Goals 10 and 11. Accurate measurement of inequality within urban areas is essential for effective mitigation strategies. Traditional methods rely on survey-based census data, which are time-consuming and delayed, while some studies use coarse proxies like nighttime lights. However, these methods are limited by resolution and fail to capture fine-grained disparities within communities. To address this, we aim to leverage accessible urban imagery, which offers detailed visual features. Two key challenges must be addressed: 1) accurately perceiving micro-level inequalities within neighborhoods, and 2) ensuring that this perception is interpretable for policy guidance. To address these gaps, we propose UI-CoT, a framework that leverages the power of urban imagery-based visual language models in urban inequality perceiving, enhanced by Chain-of-Thought prompting to improve reasoning capabilities. We fine-tune a visual language model to predict three essential neighborhood inequality indicators: the income Gini coefficient, dominant race, and racial income ratio. Extensive experiments show that our model can effectively perceive micro-level inequalities, with the incorporation of Chain-of-Thought reasoning further improving the model's performance by 17.2%. This research offers valuable insights into addressing inequalities within urban environments and demonstrates the potential of web resources in empowering urban sustainable development. The code and data are available at https://github.com/tsinghua-fib-lab/UI-CoT. Yunke Zhang, Ruolong Ma, Xin Zhang 0106, Yong Li 0008 |
WWW | 3 |
| 2024 | UV-SAM: Adapting Segment Anything Model for Urban Village IdentificationabstractUrban villages, defined as informal residential areas in or around urban centers, are characterized by inadequate infrastructures and poor living conditions, closely related to the Sustainable Development Goals (SDGs) on poverty, adequate housing, and sustainable cities. Traditionally, governments heavily depend on field survey methods to monitor the urban villages, which however are time-consuming, labor-intensive, and possibly delayed. Thanks to widely available and timely updated satellite images, recent studies develop computer vision techniques to detect urban villages efficiently. However, existing studies either focus on simple urban village image classification or fail to provide accurate boundary information. To accurately identify urban village boundaries from satellite images, we harness the power of the vision foundation model and adapt the Segment Anything Model (SAM) to urban village segmentation, named UV-SAM. Specifically, UV-SAM first leverages a small-sized semantic segmentation model to produce mixed prompts for urban villages, including mask, bounding box, and image representations, which are then fed into SAM for fine-grained boundary identification. Extensive experimental results on two datasets in China demonstrate that UV-SAM outperforms existing baselines, and identification results over multiple years show that both the number and area of urban villages are decreasing over time, providing deeper insights into the development trends of urban villages and sheds light on the vision foundation models for sustainable cities. The dataset and codes of this study are available at https://github.com/tsinghua-fib-lab/UV-SAM. Xin Zhang 0106, Yu Liu 0016, Yuming Lin 0003, Qingmin Liao, Yong Li 0008 |
AAAI | 1 |
| 2024 | GUI: A Comprehensive Dataset of Global Urban Infrastructure Based on Geospatial Visual Foundation ModelsabstractThe substantial social and financial costs of infrastructure identification impede in-depth analyses of sustainable urban design, especially in developing countries. In this paper, we present a novel framework with interactive web visualization based on geospatial visual foundation models. Leveraging this framework, we examine the urban infrastructure information in 1,178 cities worldwide, covering 93, 088 km2 areas. Cross-validation reveals that the overall accuracy of identified infrastructure achieves 67.0%. It sheds light on the sustainable development of cities and exposes the stark inequity in urban infrastructure provision for vulnerable populations. The identified urban infrastructure dataset of this study are available at https://github.com/tsinghua-fib-lab/GUI, and the interactive web application is at https://tinyurl.com/yz7xbfy3. Zhenyu Han, Xin Zhang 0106, Yanxin Xi, Tong Xia, Yong Li 0008 |
SIGSPATIAL/GIS | 2 |
| 2024 | M3 LUC: Multi-modal Model for Urban Land-Use ClassificationabstractIdentifying urban land-use types is crucial for effective resource management, urban planning, and sustainable development. However, classifying land use is complex due to the complexity of the city and the poor data available in undeveloped areas. In this work, we present the Multi-modal Model for Land-use Classification (M3LUC). Our model is the first to leverage the advanced Vision-Language Model (VLM) to better capture urban functionality through remote sensing data and Points of Interest (POI). We have also designed specific mechanisms to robustly and extensively tackle the modality missing and conflict to enhance transferability. Experiments conducted in four major cities in China demonstrate our model's superior performance in both transfer and non-transfer tasks, revealing its potential for broader applications. Sibo Li 0001, Xin Zhang 0106, Yuming Lin 0003, Yong Li 0008 |
SIGSPATIAL/GIS | 2 |
| 2024 | Long-term Detection and Monitory of Chinese Urban Village Using Satellite Imagery
Yuming Lin 0003, Xin Zhang 0106, Yu Liu 0016, Zhenyu Han, Qingmin Liao, Yong Li 0008 |
IJCAI | 2 |
| 2023 | Knowledge-infused Contrastive Learning for Urban Imagery-based Socioeconomic PredictionabstractMonitoring sustainable development goals requires accurate and timely socioeconomic statistics, while ubiquitous and frequently-updated urban imagery in web like satellite/street view images has emerged as an important source for socioeconomic prediction. Especially, recent studies turn to self-supervised contrastive learning with manually designed similarity metrics for urban imagery representation learning and further socioeconomic prediction, which however suffers from effectiveness and robustness issues. To address such issues, in this paper, we propose a Knowledge-infused Contrastive Learning (KnowCL) model for urban imagery-based socioeconomic prediction. Specifically, we firstly introduce knowledge graph (KG) to effectively model the urban knowledge in spatiality, mobility, etc., and then build neural network based encoders to learn representations of an urban image in associated semantic and visual spaces, respectively. Finally, we design a cross-modality based contrastive learning framework with a novel image-KG contrastive loss, which maximizes the mutual information between semantic and visual representations for knowledge infusion. Extensive experiments of applying the learnt visual representations for socioeconomic prediction on three datasets demonstrate the superior performance of KnowCL with over 30% improvements on R2 compared with baselines. Especially, our proposed KnowCL model can apply to both satellite and street imagery with both effectiveness and transferability achieved, which provides insights into urban imagery-based socioeconomic prediction. Yu Liu 0016, Xin Zhang 0106, Jingtao Ding, Yanxin Xi, Yong Li 0008 |
WWW | 2 |