EDBT 2026 Demo / reviewers in the wild / expert
Xixuan Hao
dblp:354/0596
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0003-0728-1944ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Traffic-R1: Reinforced LLMs Bring Human-Like Reasoning to Traffic Signal Control SystemsabstractXingchen Zou, Yuhao Yang, Zheng Chen, Xixuan Hao, Yiqi Chen, Chao Huang, Yuxuan Liang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xingchen Zou, Yuhao Yang 0002, Xixuan Hao, Chao Huang 0001, Yuxuan Liang 0002 |
ACL (1) | 4 |
| 2026 | AgentSense: LLMs Empower Generalizable and Explainable Web-Based Participatory Urban Sensing
Xusen Guo, Mingxing Peng, Xixuan Hao, Xingchen Zou, Qiongyan Wang, Sijie Ruan, Yuxuan Liang 0002 |
WWW | 3 |
| 2026 | Enhancing Ride-Hailing Forecasting at DiDi with Multi-View Geospatial Representation Learning from the Web
Xixuan Hao, Guicheng Li, Daiqiang Wu, Xusen Guo, Yumeng Zhu, Zhichao Zou, Peng Zhen 0001, Yao Yao 0004, Yuxuan Liang 0002 |
WWW | 1 |
| 2025 | UrbanVLP: Multi-Granularity Vision-Language Pretraining for Urban Socioeconomic Indicator PredictionabstractUrban socioeconomic indicator prediction aims to infer various metrics related to sustainable development in diverse urban landscapes using data-driven methods. However, prevalent pretrained models, particularly those reliant on satellite imagery, face dual challenges. Firstly, concentrating solely on macro-level patterns from satellite data may introduce bias, lacking nuanced details at micro levels, such as architectural details at a place. Secondly, the text generated by the precursor work UrbanCLIP, which fully utilizes the extensive knowledge of LLMs, frequently exhibits issues such as hallucination and homogenization, resulting in a lack of reliable quality. In response to these issues, we devise a novel framework entitled UrbanVLP based on Vision-Language Pretraining. Our UrbanVLP seamlessly integrates multi-granularity information from both macro (satellite) and micro (street-view) levels, overcoming the limitations of prior pretrained models. Moreover, it introduces automatic text generation and calibration, providing a robust guarantee for producing high-quality text descriptions of urban imagery. Rigorous experiments conducted across six socioeconomic indicator prediction tasks underscore its superior performance. Xixuan Hao, Wei Chen 0070, Siru Zhong, Kun Wang 0056, Qingsong Wen, Yuxuan Liang 0002 |
AAAI | 1 |
| 2025 | Space-aware Socioeconomic Indicator Inference with Heterogeneous GraphsabstractRegional socioeconomic indicators are critical across various domains, yet their acquisition can be costly. Inferring global socioeconomic indicators from a limited number of regional samples is essential for enhancing management and sustainability in urban areas and human settlements. Current inference methods typically rely on spatial interpolation based on the assumption of spatial continuity, which does not adequately address the complex variations present within regional spaces. In this paper, we present GeoHG, the first space-aware socioeconomic indicator inference method that utilizes a heterogeneous graph-based structure to represent geospace for non-continuous inference. Extensive experiments demonstrate the effectiveness of GeoHG in comparison to existing methods, achieving an R2 score exceeding 0.8 under extreme data scarcity with a masked ratio of 95%. The code and data are available at https://github.com/CityMind-Lab/GeoHG. Xingchen Zou, Jiani Huang 0001, Xixuan Hao, Yuhao Yang 0002, Haomin Wen, Chao Huang 0001, Chao Chen 0004, Yuxuan Liang 0002 |
SIGSPATIAL/GIS | 3 |
| 2025 | Multimodal Learning for Spatio-Temporal Data MiningabstractSpatio-temporal data mining (STDM) has become crucial in multimedia, driven by the surge of multimodal data from remote sensing, IoT sensors, social media, surveillance systems, mobile devices, and crowdsourced platforms. Traditional single-modal methods, though successful, struggle to capture real-world complexity. Integrating multiple modalities yields richer, more accurate insights, boosting spatio-temporal analysis. This half-day tutorial, MM4ST: Multimodal Learning for STDM, offers a comprehensive overview, covering STDM fundamentals, challenges in aligning and fusing heterogeneous data, advanced multimodal modeling techniques, and emerging research directions. Attendees will acquire practical knowledge to develop scalable and robust spatio-temporal mining solutions. All materials will be publicly available online. Siru Zhong, Xixuan Hao, Hao Miao 0001, Yan Zhao 0008, Qingsong Wen, Roger Zimmermann, Yuxuan Liang 0002 |
ACM Multimedia | 2 |
| 2025 | Nature Makes No Leaps: Building Continuous Location Embeddings with Satellite Imagery from the WebabstractBuilding location embedding from web-sourced satellite imagery has emerged as an enduring research focus in web mining.However, most existing methods are inherently constrained by their reliance on discrete, sparse sampling strategies, failing to capture the essential spatial continuity of geographic spaces.Moreover, the presence of confounding factors in satellite images can distort the perception of actual objects, leading to semantic discontinuity in the embeddings.In this work, we propose SatCLE, a novel framework for Continuous Location Embeddings leveraging Satellite imagery.Specifically, to address the out-of-sample query challenge of spatial continuity, we propose a geospatial refinement strategy comprising stochastic perturbation continuity expansion and graph propagation fusion, which transforms discrete geospatial coordinates into a continuous space.To mitigate the effects of confounders on semantic continuity, we introduce causal refinement, integrating causal theory to localize and eliminate spurious correlations arising from the environmental context.Through extensive experiments, SatCLE shows state-of-the-art performance, exhibiting superior spatial coherence and semantic fidelity across diverse geospatial tasks.The source code is available at https://github.com/CityMind-Lab/SatCLE. Xixuan Hao, Wei Chen 0070, Xingchen Zou, Yuxuan Liang 0002 |
WWW | 1 |
| 2024 | UrbanCross: Enhancing Satellite Image-Text Retrieval with Cross-Domain AdaptationabstractUrbanization challenges underscore the necessity for effective satellite image-text retrieval methods to swiftly access specific information enriched with geographic semantics for urban applications. However, existing methods often overlook significant domain gaps across diverse urban landscapes, primarily focusing on enhancing retrieval performance within single domains. To tackle this issue, we present UrbanCross, a new framework for cross-domain satellite image-text retrieval. UrbanCross leverages a high-quality, cross-domain dataset enriched with extensive geo-tags from three countries to highlight domain diversity. It employs the Large Multimodal Model (LMM) for textual refinement and the Segment Anything Model (SAM) for visual augmentation, achieving fine-grained alignment of images, segments and texts, yielding a 10% improvement in retrieval performance. Additionally, UrbanCross incorporates an adaptive curriculum-based source sampler and a weighted adversarial cross-domain fine-tuning module, progressively enhancing adaptability across various domains. Extensive experiments confirm UrbanCross's superior efficiency in retrieval and adaptation to new urban environments, demonstrating an average performance increase of 15% over its version without domain adaptation mechanisms, effectively bridging the domain gap. Our code and dataset are publicly accessible at https://github.com/siruzhong/UrbanCross. Siru Zhong, Xixuan Hao, Ying Zhang 0047, Yangqiu Song, Yuxuan Liang 0002 |
ACM Multimedia | 2 |
| 2024 | Terra: A Multimodal Spatio-Temporal Dataset Spanning the EarthabstractSince the inception of our planet, the meteorological environment, as reflected through spatio-temporal data, has always been a fundamental factor influencing human life, socio-economic progress, and ecological conservation. A comprehensive exploration of this data is thus imperative to gain a deeper understanding and more accurate forecasting of these environmental shifts. Despite the success of deep learning techniques within the realm of spatio-temporal data and earth science, existing public datasets are beset with limitations in terms of spatial scale, temporal coverage, and reliance on limited time series data. These constraints hinder their optimal utilization in practical applications. To address these issues, we introduce Terra, a multimodal spatio-temporal dataset spanning the earth. This dataset encompasses hourly time series data from 6,480,000 grid areas worldwide over the past 45 years, while also incorporating multimodal spatial supplementary information including geo-images and explanatory text. Through a detailed data analysis and evaluation of existing deep learning models within earth sciences, utilizing our constructed dataset. we aim to provide valuable opportunities for enhancing future research in spatio-temporal data mining, thereby advancing towards more spatio-temporal general intelligence. Our source code and data can be accessed at https://github.com/CityMind-Lab/NeurIPS24-Terra. Wei Chen 0070, Xixuan Hao, Yuxuan Liang 0002 |
NeurIPS | 2 |
| 2023 | Deformation Robust Text Spotting with Geometric PriorabstractThe goal of text spotting is to perform text detection and recognition simultaneously. Although the diversity of luminosity and orientation in scene texts has been widely studied, the font diversity and shape variance of the same character are ignored in recent works, since most characters in natural images are rendered in standard fonts. To solve this problem, we present a Chinese Artistic Dataset, termed as ARText, which contains 33, 000 artistic images with rich shape deformation and font diversity. Based on this database, we develop a deformation robust text spotting method (DR TextSpotter) to solve the recognition problem of complex deformation of characters in different fonts. Specifically, we propose a geometric prior module to highlight the important features based on the unsupervised landmark detection sub-network. A graph convolution network is further constructed to fuse the character features and landmark features, and then performs semantic reasoning to enhance the discrimination for different characters. The experiments are conducted on ARText and IC19-ReCTS datasets. Our results demonstrate the effectiveness of our proposed method. The datasets and models will become publicly available after publication. Xixuan Hao, Aozhong Zhang, Xianze Meng |
ICIP | 1 |
| 2023 | Relation-enhanced DETR for Component Detection in Graphic Design Reverse EngineeringabstractIt is a common practice for designers to create digital prototypes from a mock-up/screenshot. Reverse engineering graphic design by detecting its components (e.g., text, icon, button) helps expedite this process. This paper first conducts a statistical analysis to emphasize the importance of relations in graphic layouts, which further motivates us to incorporate relation modeling into component detection. Built on the current state-of-the-art DETR (DEtection TRansformer), we introduce a learnable relation matrix to model class correlations. Specifically, the matrix will be added in the DETR decoder to update the query-to-query self-attention. Experiment results on three public datasets show that our approach achieves better performance than several strong baselines. We further visualize the learnt relation matrix and observe some reasonable patterns. Moreover, we show an application of component detection where we leverage the detection outputs as augmented training data for layout generation, which achieves promising results. Xixuan Hao, Danqing Huang, Jieru Lin, Chin-Yew Lin |
IJCAI | 1 |